OpenAI’s GPT-6.1 Sol matches flagship Astra on DeepSWE v1.1 at $0.65 per task—a 7x cost reduction—while cutting autonomous computer-use safety violations from 17.4% to 4.3% and dropping cached input pricing to $0.10/1M tokens.
OrcaSAQ-2 shrinks a 27B coding model to about 12GB. We examine its fidelity results, agent benchmarks, and the practical limits of running it on a 16GB GPU.
GPT-6 Luna is free on Freebuff for eligible users. Here’s the five-hour catch, how to check your account, start coding, and protect private code.
Pixel Canary’s 90% pass@4 benchmark met a different reality in Cline: two runs, 38 minutes of wall time, socket errors, and no code to test.
Whiteboard AST diffs, Cursor Rollouts production bots, and Gemini CLI 0.61 safeguards ignite a post-IDE agent revolt as DHH declares pencils down.
Google Antigravity SDK now runs local models offline. Use this interactive CLI trick to run Ollama, LM Studio, or Gemma 4 LiteRT on your local machine.
Gemini 3.8 Flash takes on GPT-6 Luna. Compare 73.7% vs 66.6% DeepSWE, $0.10/M token pricing, TPU v6e vs Astra routing, and 2.8% deception rates.
Xiaomi streamed its $3.5M real-time RL run for MiMo-V2.6. An architectural audit of live cluster telemetry, fatal OOM restarts, GRPO mechanics, and benchmark tops.
OpenAI drops GPT-6 Sol at $2/M input with 50% fewer errors, hitting 68.8% on DeepSWE v1.1 to match Claude Fable 5 at 80% lower cost. Full tech breakdown.
OpenAI releases GPT-6 Luna at $0.10/M input and $0.50/M output, scoring 66.6% on DeepSWE to match Claude Opus 5 at 93% lower per-task operational cost.