一个名为Livenerf的开源基准项目已启动,用于监测Claude Opus 5.5模型在发布后是否存在性能下降1。Claude Opus 5.5于2026年9月22日发布1,该项目从2026年9月24日22:10 UTC开始,计划运行30天以检测模型精度随时间的漂移情况1。
该项目采用2,336个精选问题的面板进行测试,包括GPQA Diamond、MMLU-Pro、竞赛数学和AIME 2025-26等基准1。其中78个问题构成主要测试面板,Claude Opus 5.5在首次尝试时的准确率约为93%1。项目将前10天作为基线,后两个10天窗口用于检测潜在变化,可检测的精度变化阈值约为每10天窗口7.5个百分点1。
Livenerf采用多项严谨的检测标准,包括成对逐项评分差异、聚类标准误和输出token计数等关键指标1。确认性能下降需要在两个连续10天窗口中的99%置信区间排除零、效应至少达到3个百分点,且控制组无相同移动1。截至2026年9月29日,该项目已收集6天数据,6个测试日均完成了完整90个样本的运行1。所有测试结果和方法完全透明公开1。
Livenerf is an open-source benchmarking project designed to monitor whether Claude Opus 5.5 experiences performance degradation following its release 1. The initiative launched on September 24, 2026, and will run continuous deterministic tests over a 30-day period 1. The project uses a panel of 2,336 carefully selected questions to detect whether model accuracy drifts over time, while also tracking output token count as a supplementary signal, with all results and methodologies fully transparent 1.
Claude Opus 5.5 was released on September 22, 2026 1. The testing framework divides the monitoring period into three phases: the first 10 days establish a baseline, followed by two consecutive 10-day windows to detect changes 1. The primary test panel comprises 78 questions drawn from benchmarks including GPQA Diamond, MMLU-Pro, competition mathematics, and AIME 2025-26 1. The model demonstrates an accuracy rate of approximately 93% on its first attempt 1. The project can detect precision changes of roughly 7.5 percentage points within a 10-day window 1. Detection requires a 99% confidence interval excluding zero across two consecutive 10-day windows, with an effect size of at least 3 percentage points and no equivalent movement in control groups 1.
As of September 29, 2026, the project had collected data from 6 of its planned 30 days, with all six test days completing full runs of 90 samples 1.
评论
还没有评论,欢迎留下第一条。