Ilya Gusev、Matthias Petri 和 Andrey Styskin 发布了 NEEDLE,这是一个面向 AI Agent 的实时开源搜索引擎质量基准测试 1。NEEDLE 代表新闻、日常、专家、长尾和法律评估(News, Everyday, Expert, Deep-tail, and Legal Evaluation) 1。该基准旨在解决现有静态基准易被过拟合和数据泄露的问题 1。通过持续更新的查询,NEEDLE 评估了搜索引擎在新闻、金融、学术、长尾实体和法律等场景下的表现 1。
该基准的查询部分来自真实 Agent 搜索日志,部分为生成数据 1。SIGIR'26 研究分析了 1400 万次 Agent 搜索请求,显示 Agent 的拼写错误率约为人类的三分之一 1。该基准还对比了 Keenable 与 Google、Bing、Brave 及 Tavily 等 AI 搜索初创公司的表现 1。Keenable 的 p50 延迟为 200ms,目标 p95 延迟为 200ms 1。Keenable 六个月前建立首个索引,四个月前上线新闻覆盖,三个月前上线查询语法,两个月前上线摘要功能 1。
Ilya Gusev, Matthias Petri, and Andrey Styskin have released NEEDLE, a real-time, open-source benchmark designed to evaluate search engine quality for AI agents 1. NEEDLE stands for News, Everyday, Expert, Deep-tail, and Legal Evaluation 1. The benchmark aims to overcome the overfitting and data leakage issues inherent in existing static benchmarks by utilizing continuously updated queries to assess search engine performance across diverse scenarios, including news, finance, academia, long-tail entities, and legal domains 1.
The queries used in the NEEDLE benchmark are sourced from both real agent search logs and generated data 1. The benchmark allows for direct performance comparisons between Keenable and established platforms like Google, Bing, and Brave, alongside AI search startups such as Tavily 1. Keenable has recorded a p50 latency of 200ms and set a target p95 latency of 200ms 1. Keenable built its first index six months ago, launched news coverage four months ago, introduced query syntax three months ago, and added summarization features two months ago 1.
A study presented at SIGIR'26 analyzed 14 million agent search requests and found that agents exhibit a typo rate of approximately one-third that of humans 1.
评论
还没有评论,欢迎留下第一条。