OpenAI drops GPT-6 Sol at $2/M input with 50% fewer errors, hitting 68.8% on DeepSWE v1.1 to match Claude Fable 5 at 80% lower cost. Full tech breakdown.
OpenAI releases GPT-6 Luna at $0.10/M input and $0.50/M output, scoring 66.6% on DeepSWE to match Claude Opus 5 at 93% lower per-task operational cost.
Extending test-time compute can collapse LLM accuracy. Here is the forensic math on Goodhart gaming, attention drift, and why latent recurrent depth wins.
Technical benchmark audit of DeepSWE v1.1 by Aditi Sharma. Debunking git reflog leaks and test tampering, while exposing how 166-turn Test-Time Compute subsidizes 74% Pass@1 rates.
Why running autonomous coding agents on GPT-6 Astra leads to 19-minute execution freezes, and how leaked alpha benchmarks of GPT-6 Sol running 6.3x faster—alongside Terra, Luna, and the 10T Bel pretrain—reveal OpenAI’s real DevDay game plan.