记账、组合,还是不可达的黄金?MemoryAgentBench 冲突解决分数与冻结 Last-Write Resolver 的对比解读
Bookkeeping, Composition, or Unreachable Gold? Reading MemoryAgentBench's Conflict-Resolution Scores Against a Frozen Last-Write Resolver
Egor Pakhomov · Erik Nijkamp
中文摘要
MemoryAgentBench 的冲突解决(Conflict Resolution)分项通常被解读为衡量「选择性遗忘」。将基准自身规则——关于某事实的最新陈述获胜——作为零学习 resolver 实施,并冻结在四份事实列表之一上。按官方指标,该规则可回答 80.25% 的问题(在三个留出列表上为 74.5%)。在其余问题中,67 题存在已发布的黄金答案,Last-Write 图无法到达,但被覆盖的陈述可以(例:「The capital of India is New Delhi.」被「The capital of India is Grosseto.」覆盖;黄金答案为 New Delhi);此类题目在 262K 数据规模下占多跳问题的三分之一。两个长上下文模型以及基准 BM25 agent 的预注册近似重实现(每题仅保留一次运行与最终结果)在规则可解的题目上分别得分 84.7%、82.6% 与 41.6%,而在该 67 题上分别仅得 10.4%、11.9% 与 6.0%。失败部分源于可达性分项,并附带少量解析器作用域残余;在 per-item 分项而非总分上,才能解读该基准的得分。
关键要点
- 01MemoryAgentBench 冲突解决分项被广泛解读为衡量选择性遗忘,本文以零学习 last-write 规则作为冻结 resolver 进行重读
- 02该规则即可答对 80.25% 问题(三份留出列表 74.5%),说明大量所谓「冲突」实为记账问题
- 03其中 67 题在 262K 规模下占多跳题三分之一,其黄金答案被覆盖陈述遮蔽,Last-Write 图本身无法到达
- 04两个长上下文模型与 BM25 agent 重实现在该 67 题上骤降至 10.4%/11.9%/6.0%,分数应按 per-item 分项而非总分解读
- 05失败主要源自可达性分项,另有少量解析器作用域残余
解读
尚无解读。
原始英文摘要
arXiv:2610.09193v1 Announce Type: cross Abstract: MemoryAgentBench's Conflict Resolution split is read as measuring "selective forgetting". We execute the benchmark's own rule - the newest statement about a fact wins - as a zero-learning resolver frozen on one of the four fact lists. Under the official metric the rule answers 80.25% of the questions (74.5% on the three held-out lists). Of the rest, 67 items have a released gold that the last-write graph cannot reach but overwritten statements would ("The capital of India is New Delhi." superseded by "The capital of India is Grosseto."; gold New Delhi); such items are a third of the multi-hop questions at 262K. Two long-context models and our pre-registered approximate re-implementation of the benchmark's BM25 agent, one retained run per item and outcomes only, score 84.7%, 82.6% and 41.6% on the items the rule solves against 10.4%, 11.9% and 6.0% on those 67. The failures are a reachability split plus a small parser-scope residual; the per-item split, not the aggregate, is the unit at which a score here can be read.