开发者发布了Jeff,一套基于Qwen和Gemma模型微调的轻量级决策模型系列1。该项目提供0.8B和2B两个版本,采用与Jev兼容的请求格式1,可应用于队列路由、用户意图分类、审核标签等场景1。
性能方面,Jeff-Qwen3.5-0.8B版本在RTX PRO 6000工作站GPU上的推理延迟约为22毫秒,在Apple M4 Max上约为28毫秒1。零样本分类准确度初始为31.7%,经语音导航任务微调后可达95.8%1。该模型在4,599个公开基准问题的测试中表现接近或超越Jev1。
训练过程完全依托本地硬件完成,仅使用单块RTX PRO 6000 GPU,无需云计算资源1。0.8B模型训练耗时约2小时,2B模型约3.5小时1。合成训练数据由开源模型Qwen3.8-Flash-Next生成1。项目基于开源的AutoJev配方开发1,代码采用MIT许可证发布,模型权重采用Apache 2.0许可证1。
A new project called Jeff delivers lightweight decision models designed for rapid zero-shot classification tasks.1 The models, built on fine-tuned versions of Qwen and Gemma, are available in 0.8B and 2B parameter sizes and achieve inference latencies around 30 milliseconds.1 Jeff implements a request format compatible with Jev, enabling applications such as queue routing, user intent classification, and content moderation.1
The 0.8B variant executes in approximately 22 milliseconds on an RTX PRO 6000 workstation GPU and 28 milliseconds on an Apple M4 Max processor.1 Training times are minimal, with the smaller model requiring roughly two hours and the larger variant approximately 3.5 hours, all completed on a single local RTX PRO 6000 GPU without reliance on cloud computing resources.1 The training data was synthetically generated using the open-source Qwen3.8-Flash-Next model.1
Performance metrics show zero-shot classification accuracy of 31.7 percent without fine-tuning, improving to 95.8 percent following voice navigation-specific tuning.1 Testing across 4,599 public benchmark questions demonstrated that Jeff's performance approaches or surpasses that of Jev.1 The project builds on the open-source AutoJev recipe; the code is distributed under the MIT license while model weights use the Apache 2.0 license.1
评论
还没有评论,欢迎留下第一条。