开发者在Hacker News上发布了JevBench,这是一个专为评估类Jev决策模型性能设计的可复现基准测试工具1。该基准包含534个英文决策任务,v1.3评分综合考虑机会校正智能度、校准度、速度和成本1。
根据当前排行榜,Jev以74.4分排名第一,其后依次为SemIf(73.1分)、djev(73.0分)、Winnow-12B Q8(71.2分)和reflex 4B(70.3分)1。该项目采用MIT许可证,提供公开项目、冻结工件和评分代码1。
该基准测试存在若干限制条件:仅支持英文;延迟测量来自一台德国服务器;本地和演示延迟进行了披露的×2调整(+150毫秒)1。
A developer has unveiled JevBench on Hacker News, a reproducible benchmark tool designed to evaluate the performance of Jev-like decision models.1 The benchmark comprises 534 English-language decision tasks that comprehensively assess model accuracy, latency, and cost.1
The v1.3 scoring methodology incorporates opportunity-adjusted accuracy, calibration, speed, and cost considerations.1 According to the current leaderboard, Jev leads with a score of 74.4, followed by SemIf at 73.1, djev at 73.0, Winnow-12B Q8 at 71.2, and reflex 4B at 70.3.1 The benchmark is released under an MIT license and provides open-source projects, frozen artifacts, and scoring code for transparency and reproducibility.1
The tool operates under several acknowledged limitations: it supports English-language tasks only, latency measurements are sourced from a single German server, and local and demonstration latency figures include a disclosed 2x adjustment with an additional 150 milliseconds buffer.1
评论
还没有评论,欢迎留下第一条。