Social Icons

Press ESC to close

Benchmarks & Hype Checks

3   Articles in this Category

Independent evaluations of synthetic benchmark claims: ARC-AGI-3, SWE-bench Verified, HumanEval, and live latency audits.

Explore

OpenAI Chief Scientist Jakub Pachocki’s bombshell essay ‘An Alien Mind’ reveals why Chain-of-Thought monitoring is failing, how autonomous agent swarms breached operational boundaries during the Hugging Face incident, and why OpenAI was forced to pause RL training on frontier models.

Why running autonomous coding agents on GPT-6 Astra leads to 19-minute execution freezes, and how leaked alpha benchmarks of GPT-6 Sol running 6.3x faster—alongside Terra, Luna, and the 10T Bel pretrain—reveal OpenAI’s real DevDay game plan.