跳到正文
elvis· @omarsar0 · X·· 4 小时前AI 评分58
AI 导读

微软及合作者提出 CorpusMap,预先解析文档集合中的重复实体,为每个实体生成一个页面并链接到所有提及它的文档,原始文档保持原位。智能体读完一篇文档后可沿实体跳转到相关文档,避免重复检索同一证据。在 7 个模型和三个基准上,答案质量提升 6.4 至 11.7 分,输入 token 减少 34% 至 57%,并优于 LLM Wiki 层和另外三种导航层;该地图无需 LLM 调用即可构建,并可在新文档到达时更新。论文地址 https://arxiv.org/abs/2609.37226。

正文

Recommended. LLM agents love structure, so it's no surprise that a corpus improves agentic search.

引用DAIR.AI@dair_ai
Banger paper from Microsoft and colleagues. If you run agents that search a large document collection, this one is worth your time. (bookmark it) They introduce CorpusMap, which resolves recurring entities across the collection in advance and gives each entity a page that links to every document that mentions it. The original documents stay in place. The agent reads a document, follows an entity to related documents, and avoids searching for the same evidence again. Across 7 models and three benchmarks, answer quality goes up 6.4 to 11.7 points while input tokens drop 34% to 57%. It also beats an LLM Wiki layer and three other navigation layers. The map can be built without LLM calls and updated as new documents arrive. Paper: https://arxiv.org/abs/2609.37226 Chat with Paper: https://academy.dair.ai/papers/follow-the-entities-a-corpus-map-for-agentic-search-2609.37226
在 X 查看被引用的帖子

来源:elvis · x.com