Microsoft has released FrogNano, a compact coding-agent checkpoint derived from Qwen3.5-4B. Its researchers report a 61.5% solve rate on SWE-bench Verified after reinforcement learning on about 1,500 synthetic software tasks. The useful engineering finding is broader than the score: in a controlled comparison before that training, changing the agent interface alone lifted the base model’s solve rate from 8.3% to 37.2%.
That harness experiment warns against treating coding ability as a property of weights alone. The same model may spend its budget navigating a repository and editing files, or spend it failing to express a valid completion action. FrogNano’s release includes both downloadable weights and the Leaf harness code, so a user can inspect both sides of that result.
A smaller interface removed a failure mode
Microsoft’s first setup used the more elaborate R2E-Gym/SWE-Agent interface. The paper says roughly 96% of Qwen3.5-4B trajectories reached the turn limit, often without submitting. Leaf gives the model five typed tools—read, write, edit, glob and bash—and treats a response without tool calls as a final answer. Holding the other rollout settings fixed, the base checkpoint moved from 8.3% to 37.2% solved on SWE-bench Verified. A larger comparison model did not show the same sensitivity in that experiment.
Leaf is a specific operating environment, not just a prompt. The model sees the issue, repository state and prior tool results, may issue one or more tool calls, then receives outputs and continues. The released evaluation harness runs tasks in isolated Kubernetes sandboxes. Reproducing the headline result therefore means matching the checkpoint, tokenizer, tool schema, server parsers, benchmark images and budgets, not merely loading a 4B model into a generic chat window.
Synthetic tasks helped, with important test boundaries
TaskPilot generated about 300 candidate tasks per iteration, used the current checkpoint to calibrate difficulty, and repeated this five times. The research report shows SWE-bench Verified rising from a 43.0% starting point in its iterative training series to 61.5% after the fifth round. That series is not a clean arithmetic continuation of the separate 8.3%-to-37.2% harness test; the comparisons answer different questions.
Microsoft also reports 37.6% on SWE-bench Pro, 31.1% on Terminal-Bench 2.0 and 47.3% on PatchEval-Verified. The latter three are held-out benchmark families, while SWE-bench Verified was used for validation during the research. The evaluations used Leaf, a 131K-token context ceiling, up to 150 agent steps, temperature 0.6 and scores averaged over three seeds. Those settings give a small model substantial room to search and retry; they are part of the measured system.
The paper’s failure analysis remains sobering. Among failed SWE-bench Verified trajectories, it classifies 90.8% as reasoning gaps and 7% as premature termination. In its test-behavior breakdown, 69.8% never fixed the issue and 24.8% introduced a regression. The common reasoning mistakes were attacking the wrong root cause or misreading the specification. A production workflow still needs review and tests that exercise requirements beyond the visible patch verifier.
Local weights do not make the full evaluation cheap
The model is roughly 4.66 billion parameters in the Hub metadata. At BF16, weights alone are roughly 9.3 GB before runtime buffers and the context cache; quantization can reduce weight storage, but it does not remove the memory and time cost of long agent traces. Microsoft’s published evaluation recipe also needs a Kubernetes cluster and benchmark containers. Running a short local coding task is a much lighter proposition than reproducing the 500-task, three-seed SWE-bench experiment.
There is one release detail worth resolving before redistribution. The Hugging Face repository’s metadata labels the model mit, while the model-card prose calls its license Apache-2.0; the separate Leaf code repository displays MIT. That discrepancy does not prevent reading or testing the public files, but a team packaging the weights should confirm the intended model license with Microsoft rather than infer it from the harness.
FrogNano’s best-supported claim is that compact coding agents can improve substantially when task generation and tool interface are designed for their failure modes. To judge it for a real repository, run the same issues through the published Leaf workflow and the intended local setup, record solved tasks, regressions, steps, memory and wall time, and keep the model comparison on an identical harness.
