OpenCode Zen上的免费隐形模型big-pickle在Scale AI的SWE Atlas Codebase QnA基准测试中达到了50.8%的任务解决率,共解决了124个任务中的63个[1]。在Mini-SWE-Agent框架类别中,该模型的表现超越了官方排行榜上的所有其他条目,仅被两个Claude原生模型超越[1]。
此次评估于2026年8月11日进行,采用了官方开源工具、公开任务数据和指定的评判模型(claude-opus-4-5-20251101)[1]。模型运行消耗了6.74亿个输入token和430万个输出token,在隐形期间免费使用[1]。根据估算,完整运行成本约为Modal计算费用70美元加上Anthropic API费用25美元[1]。评估中使用的计算资源为4个CPU和8GB内存,低于官方声明的16个CPU和16GB内存配置[1]。单次试验的标准误差为±4.5个百分点[1]。
Big-pickle, a free stealth model available on OpenCode Zen, has demonstrated strong performance on Scale AI's SWE Atlas Codebase QnA benchmark, resolving 63 out of 124 tasks for a 50.8% success rate [1]. The evaluation, conducted on August 11, 2026, employed official open-source tools, publicly available task data, and Claude Opus 4.5 as the designated evaluation model [1].
Within the Mini-SWE-Agent framework category, big-pickle surpassed all entries currently listed on the official leaderboard, with only two native Claude models ranking higher [1]. The assessment consumed 674 million input tokens and 4.3 million output tokens, with the model incurring no cost during its stealth period [1]. The evaluation infrastructure utilized 4 CPU cores and 8 GB of memory—below the declared requirements of 16 CPU cores and 16 GB—while the complete operational cost was estimated at approximately $70 for Modal computation and $25 for Anthropic API access [1]. The single trial standard error was measured at ±4.5 percentage points [1].