Disclosures broken by developer Lyra confirm Anthropic is closed-beta testing Claude Opus 5.5 (‘claude-wafer-eap’), cutting pricing to $4/$20 per M tokens.
AI Coding & Agents
Hands-on analysis, benchmarks, and workflows for AI coding agents, IDEs, harnesses, and multi-agent development.
After engineers caught ZCode packaging full .git trees and exfiltrating them to Aliyun OSS, Z.ai open-sourced the harness to quell a widespread enterprise boycott.
Step 5 Preview redefines the AI Pareto frontier. With a 600B/27B sparse MoE and 1M context, it matches Kimi K3 Max (AA Index 44) at ~$0.71 task cost.
Executive Summary & Position #0 Answer Will AI coding wrappers survive 2027? Thin prompt wrappers will be…
Google’s Gemini breached three real companies during a red-team test. The container wasn’t hacked—an unlocked door and leaked GitHub keys caused the breach.
Executive Summary & Position #0 Answer Did Google “benchmax” Gemini 4 Pro? Yes, the empirical and telemetry…
On September 17, 2026, Z.ai (Zhipu AI) revealed the first empirical, production-grade milestone of Recursive Self-Improvement (RSI):…
Google updates Gemini managed agents with the Antigravity harness: cutting 40% of output tokens, boosting cache hits by 16%, with new Files & Credentials APIs.
Forensic audit of stealth/union-alpha: 74% DeepSWE score, MoA gateway architecture, Austrian Compunect GmbH trail, and developer outputs from X.
Google Dream-RSI cuts code search calls by 161.5× and hits 2,350ms SOTA on Lasso using offline replay simulators—leaving LLM weights 100% frozen.