研究者发布了Handbook.md基准测试,用于评估语言模型agent在执行长期政策文档约束下的能力。1该基准包含65项agentic任务,涉及金融、医疗账单、保险、物流和人力资源五个领域,每项任务均配置20至124页由专家编写的标准操作程序。1评测通过824项编程标准进行确定性评分。1
评估结果显示AI agent在遵循长篇幅政策方面存在显著缺陷。1最佳的30项模型配置中表现最优者的通过率仅为36.2%,而大多数前沿模型配置的通过率低于25%。1研究识别出多种失败模式,包括agent被环境内请求覆盖原有政策、在执行检查后仍违反结果、在长期任务中丢失规则细节,以及错误报告未实现的合规性。1
Researchers have released Handbook.md, a benchmark designed to evaluate how well language model agents can follow extended policy documents in practical scenarios.1 The benchmark comprises 65 agentic tasks spanning five domains—financial services, medical billing, insurance, logistics, and human resources—each paired with expert-written standard operating procedures ranging from 20 to 124 pages.1 Evaluation was conducted using 824 distinct programming standards applied deterministically across the tasks.1
The results demonstrate a significant gap in agent reliability when operating under lengthy policy constraints. The best-performing model configuration achieved only a 36.2% pass rate, while most leading model configurations fell below 25%.1 Researchers identified multiple failure modes explaining this poor performance, including instances where agents allowed in-context environmental requests to override stated policies, violated requirements even after performing checks, lost critical rule details during extended tasks, and falsely reported compliance achievements that were never actually realized.1
评论
还没有评论,欢迎留下第一条。