ExperienceIndex:基于语料的构件经验记忆
ExperienceIndex: Artifact-Grounded Memory
Peter Baile Chen · Geoffrey X. Yu · Xinming Liu · Samuel Madden · Dan Roth · Jacob Andreas · et al.
中文摘要
知识密集型任务需要基于共享构件语料(如判例或科学文献)进行多问推理。人类在反复接触这些语料时,会自然积累关于构件的隐性经验,从而在新任务中快速锁定全部相关构件。然而现有 AI 智能体缺乏合适的记忆机制来构建或复用这种构件级经验,导致答案质量下降且在线成本偏高。已有记忆方案主要从历史求解轨迹中提取并复用信息,但侧重于用户偏好、事实属性或抽象推理模式,而非持续性的构件专属知识。本文提出 ExperienceIndex,作为 AI 智能体的全新经验层,基于先验推理轨迹捕获与复用构件知识。它存储两种互补的经验形式:(i) 单构件经验,总结某构件在过往任务中的贡献;(ii) 构件对经验,编码推理过程中发现的构件间结构关系。ExperienceIndex 以轻量中间件形式集成,通过经验检索机制引导智能体找全新任务的相关构件,同时提升答案质量与效率。在多类语料和不同搜索框架的智能体方案上,ExperienceIndex 一致带来收益,答案质量最高提升 11.0 个点,在线美元成本最高降低 50.5%。还展示了两项额外优势:(i) 跨任务泛化——从 text-to-SQL 任务积累的经验可迁移到同一构件语料上的事实问答任务;(ii) 师生学习——强模型积累的经验可使弱模型达到相近性能。
关键要点
- 01问题:知识密集型任务中,现有 AI 智能体的记忆机制聚焦于用户偏好与抽象推理模式,缺乏对构件级持续经验的捕获与复用能力,造成答案质量与效率损失。
- 02方法:ExperienceIndex 以轻量中间件形式存储两类经验——单构件经验与构件对经验,并通过经验检索机制在推理时引导智能体定位相关构件。
- 03结果:在多类语料与不同搜索框架下,答案质量最高提升 11.0 个点,在线美元成本最高降低 50.5%,并验证了跨任务泛化与师生迁移两种迁移收益。
- 04局限:摘要未披露不同语料与基线下的统计显著性及细粒度失败案例,经验检索对构件规模与异构性的可扩展性边界也未深入讨论。
解读
尚无解读。
原始英文摘要
arXiv:2610.10091v1 Announce Type: new Abstract: Knowledge-intensive tasks require answering many questions by reasoning about a shared corpus of artifacts (e.g., court cases, or scientific literature). As humans interact with these corpora, they naturally accumulate experiential knowledge about artifacts, enabling them to quickly identify the complete set of relevant artifacts for each new task. However, existing AI agents lack appropriate memory solutions to build or reuse such artifact-grounded experience, leading to lower answer quality and higher online cost. Existing memory solutions extract and reuse information from prior task-solving traces, but they primarily focus on user preferences, factual attributes, or abstract reasoning patterns rather than persistent artifact-specific knowledge. We introduce ExperienceIndex, a novel experience layer for AI agents that captures and reuses knowledge about artifacts based on prior reasoning traces. ExperienceIndex stores two complementary forms of experience: (i) single-artifact experiences that summarize an artifact's contribution to prior tasks and (ii) artifact-pair experiences that encode structural relationships discovered during past reasoning. Integrated as lightweight middleware, ExperienceIndex uses an experience retrieval mechanism to guide agents toward the complete set of relevant artifacts for new tasks, improving both answer quality and efficiency. Across diverse corpora and agentic solutions with different search frameworks, ExperienceIndex delivers consistent gains, raising answer quality by up to 11.0 points and reducing online dollar cost by up to 50.5%. We further demonstrate two benefits: (i) cross-task generalization, where experiences accumulated from text-to-SQL tasks transfer to factoid QA tasks over the same artifact corpus, and (ii) teacher-student learning, where experiences from a stronger model enable a weaker model to reach comparable performance.
同方向论文 · cs.CL
查看全部 →EngramEdit:通过条件记忆实现大语言模型的解耦知识更新
2610.10533Prompt 应该做更多:检索指令对嵌入模型(embedding models)的影响
2610.10508PHRBench:面向大语言模型幻觉后推理(后幻觉推理)的行为评测
2610.10455CoTrace:通过 Harness-Model 协同演化训练终端 Agent 的数据配方
2610.10426使用大语言模型实现爱沙尼亚语文档级文本简化
2610.10378面向任务进度的行动学习:从紧凑教师监督中蒸馏小型智能体
2610.10332