OrcaSAQ-2 27B compresses Qwen3.8-27B from about 54GB to a 12.3GB checkpoint. OrcaRouter reports 93.2% agreement with the original model’s top token choices, but the checkpoint needs more than 12GB of GPU memory to serve. Its own deployment guidance targets a 16GB card and roughly 32,000 tokens of practical interactive context. Source: OrcaRouter model card.
A coding agent can take dozens of steps to fix one bug. It reads files, edits code, runs a command, studies the error and tries again. If the model makes a poor decision halfway through, the next steps may be spent recovering from that decision. That is the problem OrcaRouter says it built OrcaSAQ-2 27B to address.
Released on September 24, 2026, OrcaSAQ-2 is a compressed version of Qwen3.8-27B aimed at coding, terminal work and agents that use several tools over long sessions. The weights are available under Apache 2.0. OrcaRouter’s announcement on X gives a slightly different storage figure of 12.06GB; the current model card lists 12.3GB. Both describe a checkpoint of roughly 12GB, rather than the total GPU memory needed to run it.
Compression is the achievement; preserved decisions are the test
OrcaSAQ-2 is a quantized checkpoint, not a newly trained 27B model. Quantization stores weights with fewer bits. OrcaRouter calls its approach sensitivity-aware mixed precision: parts of the model that are more sensitive to compression receive more precision. The decoder averages 3.21 bits per weight, and the published checkpoint is about 77% smaller than the BF16 original. OrcaRouter has released the weights and serving integration, but has not disclosed its detailed calibration or precision-allocation method. Source: OrcaRouter model card.
To measure what survived, OrcaRouter compared the compressed weights with Qwen3.8-27B BF16 on WikiText-2. It reports that perplexity moved from 5.6468 to 5.6482, an increase of about 0.02%. The models selected the same highest-ranked next token at 93.2% of positions in that evaluation. Mean divergence between their next-token probability distributions was reported as 0.031. Source: OrcaRouter’s evaluation table.
That 93.2% figure measures token choices under the evaluation conditions. It does not mean the two models deliver the same complete answer 93.2% of the time. A changed token can redirect the rest of a response. In an agent, it could alter a tool call, a file edit or the command selected next. OrcaRouter explicitly says its perplexity result does not guarantee identical downstream performance.
The agent benchmark scores need a controlled comparison
OrcaRouter reports 70.0% on SWE-bench Verified and 58.4% on Terminal-Bench 2.1. Those are notable claimed results for a checkpoint this small. The comparison table, however, places them alongside scores obtained with different models and agent setups. Tools, time limits, reasoning budgets and the software surrounding a model can all affect completion rates. OrcaRouter cautions against reading the table as a strict model ranking. Source: OrcaRouter benchmark notes.
The decisive experiment would run OrcaSAQ-2 and the original Qwen3.8-27B through the same agent, on the same tasks, with the same settings. That would reveal where compression changes completed work rather than just next-token predictions. The published model card does not provide that controlled agent comparison.
Why a 12GB checkpoint calls for a 16GB GPU
The weight file is only one part of inference memory. vLLM, the working memory used for context, concurrent requests and optional multi-token prediction also consume GPU capacity. OrcaRouter targets a 16GB GPU and suggests starting around 32,000 tokens of interactive context on that tier. The model architecture supports up to 262,144 tokens, but that specification does not mean the full window fits into a 16GB deployment. Source: OrcaRouter deployment guidance.
Serving choices change the balance. In OrcaRouter’s model-card measurements under a 15.7GiB GPU memory cap, multi-token prediction increased single-stream generation from 65.3 to 90.1 tokens per second. It also reduced the reported key-value cache pool from 29,354 to 14,563 tokens and lowered throughput under eight or sixteen concurrent streams. Those figures belong to OrcaRouter’s measured setup; they are not a speed guarantee for every 16GB card. Source: OrcaRouter serving results.
The model requires OrcaRouter’s serving integration. It is also text-only: the vision component available in the original Qwen3.8-27B is not included in this checkpoint. These are practical considerations for anyone expecting a simple download-and-run replacement for the base model.
What the release establishes
OrcaSAQ-2 27B brings a large model’s weights into a size that can be served on a single 16GB GPU. The published fidelity measurements show that the compressed model stays close to BF16 on one next-token evaluation. Its claimed agent scores are promising, but their strongest interpretation awaits independent, same-setup tests against the original Qwen3.8-27B. For developers considering local coding agents, the immediate question is concrete: with your tools and your context length, does it finish the task reliably enough to justify the smaller memory footprint?
