The Spectrum Dispatch News

entertainment

Livenerf benchmark tracks Claude Opus 5.5 for post‑launch performance changes

The open‑source livenerf project runs a fixed set of prompts daily on Claude Opus 5.5 to detect any quiet degradation after its September 22, 2026 release; as of September 29, six 

Livenerf benchmark tracks Claude Opus 5.5 for post‑launch performance changes

The livenerf repository, hosted at https://github.com/ninjahawk/livenerf, describes a deterministic benchmark designed to answer whether a frontier model gets worse after it ships. The tool focuses on Claude Opus 5.5, which was released on 2026‑09‑22. According to the project’s README, the benchmark began its first full run on 2026‑09‑24 (day 1) and is scheduled to run once a day for 30 days, with the initial 10 days serving as a baseline. Progress reported on 2026‑09‑29 indicated that six of the planned 30 days had been completed, all of them within the baseline window, and no runs had been missed. Each of those six days executed the full set of 90 samples using the same harness hash (461391b6fce64167) and a pinned CLI version (2.1.280). The benchmark’s panel consists of 78 questions drawn from GPQA Diamond, MMLU‑Pro, competition‑math and AIME 2025‑26 that the model answers inconsistently; on fresh samples the pass rate for these items rose from 54.7 % to 62.0 %, a figure used in the project’s power calculations. Livenerf’s primary metric is the paired per‑item score difference against the baseline, with clustered standard errors that remove item difficulty. A secondary metric tracks the median output token count per sample, which the authors say drops before any measurable accuracy change when the model’s effort is reduced. The document notes that a low‑effort shift would manifest as roughly a −62 % reduction in output tokens accompanied by a −8.3 ± 4.5 point accuracy change, while a medium‑effort shift would show −26 % tokens and −4.2 ± 3.9 points. As of the reported progress, the benchmark has not observed a statistically significant deviation in either the accuracy or token‑output metrics relative to the launch‑week baseline. The validation procedure, described in docs/VALIDATION.md, has passed its pre‑registered criterion, confirming that the rig can detect known degradations before accepting a null result. The project emphasizes that all numbers are generated from logs in docs/CALIBRATION.md and that the statistical approach follows Anthropic’s “Adding Error Bars to Evals” methodology, ensuring the results are not home‑brew. For anyone wishing to replicate or extend the work, the repository provides instructions for installing the required environment (Python 3.11+, uv, a logged‑in Claude Code installation), pinning the CLI to prevent harness changes, and running the daily benchmark via a cron‑style script. The ongoing series will maintain a rolling 10‑day table showing how Opus 5.5 performs relative to its launch‑week baseline, with negative deltas indicating regressions and positive deltas indicating improvements.

Livenerf benchmark tracks Claude Opus 5.5 for post‑launch performance changes

Key facts

Sources

← All posts