Needle 3 delivers 86% tool accuracy in an 8MB–29MB binary at 4,000 tok/sec. Cactus Compute’s laddered architecture replaces generative chat with edge automation.
Qwen3.8-Flash-Next previews Qwen4 with 6B active parameters, QSA sparse attention and 1M context. We examine benchmarks, caveats and local hardware.