PostHog发布了Jeeves,一款基于Qwen3.5-9B的推理增强型决策模型1。该模型采用SFT和CISPO两阶段训练方法,并集成扩散草案器技术1,在多项基准测试中展现出显著优势。
在JevBench公开难题测试集上,Jeeves达到0.935的准确度,相比原有的Jev模型的0.866有明显提升1。在通用测试集上,Jeeves准确度为0.889,超过Kev-9B的0.822和Jev的0.8571。Jeeves支持是/否、多选和评分三种问题类型,并兼容Jev API接口1。
在实际部署方面,该模型在单个H100 GPU上的中位延迟为3.3秒(包含推理过程)1。训练过程中,Jeeves在包含19,126个来自12个公开数据集的问题上进行监督微调,其中一半的样本包含推理链1。不过在MMLU基准测试上,Jeeves的准确度为0.793,相比Jev的0.900有所下降;在MMLU-Pro上为0.739,对比Jev的0.840同样存在差距1。
Jeeves, a new reasoning-enhanced decision model based on Qwen 3.5-9B, has demonstrated significant performance improvements across multiple benchmarks, outperforming comparable Kev and Jev models 1. The system employs a two-stage training approach combining supervised fine-tuning (SFT) and CISPO methods, integrated with a diffusion drafter component 1. Jeeves achieves 0.935 accuracy on JevBench's public difficult questions, surpassing Jev's 0.866, and reaches 0.889 accuracy on the test set compared to Kev-9B's 0.822 and Jev's 0.857 1.
The model supports three question types—yes/no, multiple-choice, and rating—while maintaining API compatibility with Jev 1. Training involved 19,126 questions drawn from 12 public datasets, with half including reasoning chains, using 2 epochs of SFT across 8 GPUs over 596 steps, followed by 624 steps of CISPO scheduling ending at step 402 1. Operationally, Jeeves processes requests with a median latency of 3.3 seconds including reasoning on a single H100 GPU 1, though it shows mixed results on general knowledge benchmarks, achieving 0.793 accuracy on MMLU compared to Jev's 0.900, and 0.739 on MMLU-Pro versus Jev's 0.840 1.
评论
还没有评论,欢迎留下第一条。