Superintelligent Retrieval Agent: The Next Frontier of Information Retrieval

RecSys / IR DRSA Neural Search
检索增强型智能体正日益成为访问大型组织知识库的主要交互界面;然而,目前绝大多数此类智能体仍将检索过程视为一个“黑箱”:它们先发出试探性查询,检查返回的文本片段,再据此反复调整查询语句,直至获取有用证据为止。这种做法更类似于新手在面对一个陌生数据库时的摸索式搜索,而非专家凭借对专业术语和潜在证据分布的强先验知识所开展的高效导航。其结果是造成了不必要的多轮检索、响应延迟增加,以及查全率低下。 我们提出一种名为“超级智能检索智能体”(SuperIntelligent Retrieval Agent,简称 SIRA)的新方法,将检索中的“超级智能”定义为:能够将原本需多轮试探的探索式搜索,压缩为一次面向整个语料库、具备强判别能力的单次检索操作。SIRA 并非简单地询问“哪些词与当前查询相关”,而是进一步追问:“哪些词最有可能将目标证据从语料库中大量干扰项(corpus-level confusers)中区分出来?” 在语料库端,大语言模型(LLM)预先离线为每篇文档补充其缺失的检索相关词汇;在查询端,LLM 则预测出原始查询中遗漏的关键证据词汇;此外,系统还调用基于词频的统计信息作为工具函数,自动筛除那些在语料库中完全不存在、过于常见或难以产生有效检索区分度的候选扩展词。最终的检索步骤仅需执行一次加权 BM25 检索,将原始查询与经上述验证后的扩展词组合起来一并提交。 在涵盖十个 BEIR 基准数据集及下游问答任务的全面评估中,SIRA 均展现出显著更优的性能,不仅大幅超越了稠密检索器(dense retrievers),也明显优于当前最先进的多轮智能体式基线方法。实验结果表明:一条经过精心构造的词汇级查询——由大语言模型的认知能力引导,并辅以轻量级语料库统计信息——即可在性能上远超成本高昂得多的多轮检索策略;同时,该方法仍保持高度可解释性、无需训练、且计算高效。
Retrieval-augmented agents are increasingly the interface to large organizational knowledge bases, yet most still treat retrieval as a black box: they issue exploratory queries, inspect returned snippets, and iteratively reformulate until useful evidence emerges. This approach resembles how a newcomer searches an unfamiliar database rather than how an expert navigates it with strong priors about terminology and likely evidence, and results in unnecessary retrieval rounds, increased latency, and poor recall. We introduce \textit{SuperIntelligent Retrieval Agent} (SIRA), which defines \emph{superintelligence} in retrieval as the ability to compress multi-round exploratory search into a single corpus-discriminative retrieval action. SIRA does not merely ask what terms are relevant to the query; it asks which terms are likely to separate the desired evidence from corpus-level confusers. On the corpus side, an LLM enriches each document offline with missing search vocabulary; on the query side, it predicts evidence vocabulary omitted by the query; and document-frequency statistics as a tool call to filter proposed terms that are absent, overly common, or unlikely to create retrieval margin. The final retrieval step is a single weighted BM25 call combining the original query with the validated expansion. Across ten BEIR benchmarks and downstream question-answering tasks, SIRA achieves the significantly superior performance outperforming dense retrievers and state-of-the-art multi-round agentic baselines, demonstrating that one well-formed lexical query, guided by LLM cognition and lightweight corpus statistics, can exceed substantially more expensive multi-round search while remaining interpretable, training-free, and efficient.
许愿