面向长期 LLM 智能体的情境条件化思维策略学习
Learning Situation-Conditioned Thinking Policies for Long-Term LLM Agents
Hong Su
中文摘要
长时自主运行的智能体需要在不让显式历史记忆与 LLM 上下文无限膨胀的前提下,复用累积的推理经验。现有记忆机制主要对历史内容做检索、摘要或压缩,既未直接学习在何种情境下应激活哪类思维,也未从时间分散的经验中发现新的思维知识。论文提出情境条件化的思维记忆框架,将历史推理经验转化为轻量策略,用于预测当前情境下应思考什么,而把具体推理留给大语言模型。情境可表示时间或时空演化,而非仅限于当下状态。临时经验还会被周期性跨多个独立回合加以分析,以识别重复出现的长程规律,并将其整合为新的思维知识,进一步内化进轻量策略。实验表明,该所学策略在时间规则泛化上达到 1.000 F1,将 DeepSeek 推理 F1 从 0.789 提升至 0.868,在 30,000 条历史情境下将在线处理时延由每查询 0.3636 ms 降至 0.0382 ms,并在获得充分跨经验证据后达到 1.000 的关系发现 F1 与未来思考准确率。
关键要点
- 01问题:长程智能体复用推理经验时,既有记忆机制只做检索/摘要/压缩,无法学习何时激活何种思维,也不能从分散经验中挖掘新思维知识
- 02方法:把历史推理经验编码为情境条件化的轻量策略来预测当下应思考什么;情境涵盖时间与时空演化,详细推理交由 LLM 完成
- 03机制:周期性跨回合分析临时经验,抽取重复长程规律,合并为新思维知识并内化进轻量策略
- 04结果:时间规则泛化 F1 为 1.000,DeepSeek 推理 F1 由 0.789 升至 0.868,30,000 情境下在线时延从 0.3636 ms 降至 0.0382 ms/查询
- 05效果:充分跨经验证据后,关系发现 F1 与未来思考准确率均达到 1.000
解读
尚无解读。
原始英文摘要
arXiv:2610.09590v1 Announce Type: cross Abstract: Long-running autonomous agents must reuse accumulated reasoning experience without allowing explicit historical memory and LLM context to grow indefinitely. However, existing memory mechanisms mainly retrieve, summarize, or compress past content and do not directly learn when particular kinds of thinking should be activated or discover new thinking knowledge from temporally dispersed experiences. This paper proposes a situation-conditioned thinking memory framework that transforms historical reasoning experience into a lightweight policy for predicting what should be thought about in the current situation, while leaving detailed reasoning to a large language model. Situations may represent temporal or spatiotemporal evolution rather than only current states. Temporary experiences are also periodically analyzed across multiple independent episodes to identify repeated long-range regularities, which are consolidated into new thinking knowledge and further internalized by the lightweight policy. Experiments show that the learned policy achieves 1.000 F1 on temporal-rule generalization, improves DeepSeek reasoning F1 from 0.789 to 0.868, reduces online processing time from 0.3636 ms to 0.0382 ms per query at 30,000 historical situations, and reaches 1.000 relation-discovery F1 and future-thinking accuracy after sufficient repeated cross-experience evidence.