An architectural teardown of Apple’s A20 Pro 2nm GAAFET silicon: testing the 32-core Neural Engine, 65.2 GB/s memory bandwidth wall, copper vapor chamber thermals, and iOS Jetsam limits running 3B to 7B LLMs on-device.
Qwen3.8 Flash-Next vs GLM-5.3 Flash: the short version Qwen3.8 Flash-Next and GLM-5.3 Flash are not “small models”…