STaRK: Benchmarking LLM Retrieval on Textual and Relational Knowledge Bases

回答现实世界用户查询,例如产品搜索,通常需要从半结构化知识库或涉及非结构化(例如产品的文本描述)和结构化(例如产品实体关系)信息混合的数据库中准确检索信息。然而,以往的研究大多将文本和关系检索任务作为独立的主题进行研究。为了填补这一空白,我们开发了STARK,一个大规模的半结构化检索基准,用于文本和关系知识库。我们设计了一个新颖的流程,用于合成自然而真实的用户查询,这些查询集成了各种关系信息和复杂的文本属性,以及它们的真实答案。此外,我们还进行了严格的人类评估,以验证我们基准的质量,该基准涵盖了各种实际应用,包括产品推荐、学术论文搜索和精准医学查询。我们的基准作为一个全面的测试平台,用于评估检索系统的性能,重点是由大型语言模型(LLM)驱动的检索方法。我们的实验表明,STARK数据集对当前的检索和LLM系统提出了重大挑战,表明需要构建更能够处理文本和关系方面的检索系统。
Answering real-world user queries, such as product search, often requires accurate retrieval of information from semi-structured knowledge bases or databases that involve blend of unstructured (e.g., textual descriptions of products) and structured (e.g., entity relations of products) information. However, previous works have mostly studied textual and relational retrieval tasks as separate topics. To address the gap, we develop STARK, a large-scale Semi-structure retrieval benchmark on Textual and Relational Knowledge Bases. We design a novel pipeline to synthesize natural and realistic user queries that integrate diverse relational information and complex textual properties, as well as their ground-truth answers. Moreover, we rigorously conduct human evaluation to validate the quality of our benchmark, which covers a variety of practical applications, including product recommendations, academic paper searches, and precision medicine inquiries. Our benchmark serves as a comprehensive testbed for evaluating the performance of retrieval systems, with an emphasis on retrieval approaches driven by large language models (LLMs). Our experiments suggest that the STARK datasets present significant challenges to the current retrieval and LLM systems, indicating the demand for building more capable retrieval systems that can handle both textual and relational aspects.
许愿