Livenerf is monitoring Claude Opus 5.5 after release, but it cannot yet establish that the model has been nerfed. The project’s September 29 status reports six days of observations. Its comparison requires a baseline and later windows; the earliest scheduled decision is day 30. More consequentially, its initial validation has an important sensitivity limit.

The useful question is what a score movement would establish about the service you use. A benchmark can detect some changes and miss others. Establishing that distinction before the first alarming chart appears makes the eventual result easier to interpret.

It measures one subscription route under fixed conditions

Livenerf uses a headless Claude Code subscription session with a pinned client and explicitly selected effort. Its design document describes fixed prompts and graders. That reduces changes in the measuring setup; it does not expose all changes in the service behind that setup.

The route matters. A subscription session, an API request and a coding agent working through a real repository can face different inputs and constraints. EyesTech’s earlier examination of Claude Max’s “20x” usage claim concerns the plan’s usage denominator. Livenerf concerns measured responses. A change in a quota does not by itself prove a change in answer quality, and the two should be logged separately.

The project uses Inspect’s evaluation framework. A framework can preserve logs and implement scorers; the experiment still needs a defensible sample, comparison and decision rule. A well-recorded test can answer a narrow question very well without becoming an audit of the entire product.

Most of the panel comes from one question family

The panel contains 78 questions selected near the model’s decision boundary: 59 MMLU-Pro, 12 GPQA, four competition questions and three AIME questions. That makes MMLU-Pro roughly three quarters of the panel. Selection near the boundary can improve responsiveness, but it changes what the aggregate score represents.

Repeatedly answering these questions measures stability on this selected panel. It does not directly measure repository navigation, tool recovery, editing across files or completing an application. If a future decline is concentrated in the largest question family, the overall number may conceal that concentration. Readers should want item-level and family-level results alongside the aggregate.

“Validation passed” has a narrower meaning than it sounds

The validation report reports 1,248 graded responses and three errors. Its pass criteria include compatibility with zero for an unchanged-condition comparison and a detectable output-token change between low and high effort. Those checks help establish repeatability and responsiveness to an effort change.

The same report says the tested Opus 5 versus Opus 5.5 swap was not distinguishable at its 99% accuracy or output-token thresholds. Its accuracy difference was −3.8 percentage points with a reported 95% interval of −16.3 to +8.6 points. That interval is too wide to make the point estimate a dependable ranking.

What the initial Livenerf validation establishes
ComparisonReported resultInterpretation
Unchanged conditions (A/A)Compatible with zeroNo detected difference in this check
Low versus high effortOutput-token change detectedResponsive to the tested effort change
Opus 5 versus Opus 5.5Not distinguishable at 99% thresholdsSensitivity to this swap not established

This does not make the instrument useless. It means passing those checks does not prove sensitivity to every change a user might care about. A model replacement, an effort change and a task-specific regression need not produce the same observable signature. A quiet chart cannot exclude a degradation that the panel is poorly equipped to detect.

The decision threshold and the detectable effect differ

The project’s preregistered protocol requires a decline in two ten-day windows, a 99% interval excluding zero, a magnitude of at least three percentage points and a control that does not show the same pattern. Its stated minimum detectable effect is about 7.5 points at 80% power for the planned comparison.

Three points is the minimum decline allowed by the decision rule. It is not a promise that a three-point decline will be detected reliably. Power describes how often the test would flag a specified effect under its assumptions. A team can reasonably set a decision floor below the size it has enough observations to detect consistently, but readers need both numbers.

The control deserves similar care. If both series fall, shared infrastructure or another common change might be involved. That observation would not automatically prove a faulty instrument: both services could have changed. If only the target falls, attribution still needs investigation. Weights, routing, instructions and other service components can affect the same score.

Build a second panel around your actual work

For a team deciding whether to change models, preserve a small set of representative failures with acceptance criteria that can be checked. Include easy tasks as well as difficult ones, record the effort setting and client version, and separate correctness, latency, usage limits and tool errors. Keep the test set fixed during a comparison and version any later changes.

Run repeated trials and inspect the failed outputs. Evan Miller’s discussion of evaluation error bars is useful background for why uncertainty matters. Repeating the same task gives evidence about that task’s response distribution; it does not create new task coverage.

Livenerf makes a recurring complaint measurable. The next meaningful result will be a completed, reproducible comparison, with its uncertainty and question mix intact. Until then, neither an accusation of nerfing nor a clean bill of health follows from the available series.

Last Update: September 30, 2026