LiveMACE:演化市场中LLM智能体能力的过程感知评估
LiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving Markets
Jun Zhao · Leiming Fu · Yanbo Wen · Yiding Wang · Xuantong Liu · Yang Shu · et al.
中文摘要
仅依据结果评估智能体会掩盖产生这些结果的能力本身。该问题在演化环境中尤为突出,因为结果是智能体行为与外部条件变化之间闭环交互的反映。LiveMACEBench 是一个过程感知基准,将实时金融市场作为持续 LLM 智能体的自然演化测试床。五个前沿 LLM 在匹配的工具使用、持久记忆、规则遵循与多智能体协作配置下沿连续轨迹运行。通过实际结果与源自完整决策轨迹的机制专属诊断两方面进行评估。在 30 天实时测试中,发现显著的结果-能力差异:实际收益常与能力专属度量偏离,且相似的结果可能源自明显不同的机制调用模式。轨迹级诊断进一步暴露出各类机制各自的瓶颈,表明机制可及性、有效机制使用与下游性能并非衡量智能体能力的可互换指标。LiveMACEBench 将这一区分变得可量化,使实时市场从性能排行榜转变为智能体能力的诊断环境。
关键要点
- 01问题:仅看 outcomes 无法揭示能力差异,演化市场中结果-能力偏离尤为明显
- 02方法:LiveMACEBench 用实时金融市场评估持续 LLM 智能体,从完整决策轨迹抽取机制专属诊断
- 03结果:30 天实测发现相似结果可能源于不同机制使用模式,机制可及性、有效调用与下游性能不可互换
- 04局限:评估仅覆盖五个前沿 LLM 与 30 天窗口,诊断能力受限于当前可观测的机制类型
解读
尚无解读。
原始英文摘要
arXiv:2610.09872v1 Announce Type: cross Abstract: Evaluating agents by outcomes alone can obscure the capabilities that produce them. This problem is especially pronounced in evolving environments, where outcomes reflect a closed-loop interaction between agent behavior and changing external conditions. We introduce LiveMACEBench, a process-aware benchmark that uses live financial markets as a naturally evolving testbed for persistent LLM agents. Five frontier LLMs operate along continuous trajectories under matched Tool Use, Persistent Memory, Rule Following, and Multi-Agent Collaboration configurations. We evaluate them through both realized outcomes and mechanism-specific diagnostics derived from complete decision traces. Across 30 days of live evaluation, we find a pronounced outcome-capability gap: realized returns often diverge from capability-specific measurements, and similar outcomes can arise from markedly different patterns of mechanism use. Trace-level diagnostics further expose distinct bottlenecks across capabilities, demonstrating that mechanism access, effective mechanism use, and downstream performance are not interchangeable measures of agent capability. LiveMACEBench makes this distinction measurable, turning live markets from a performance leaderboard into a diagnostic environment for agent capability