跳到正文
返回论文列表
cs.AI提交于 已译

面向语言智能体行为科学的轨迹抽象

Trajectory Abstraction for the Science of Language Agent Behavior

Tianqiang Yan

中文摘要

arXiv:2610.09237v1 公告类型:cross 摘要:语言智能体的科学研究需要能够跨任务与模型支持假设的行为变量。该研究问题被表述为学习并检验一个轨迹抽象层级。具体的递归流程先测量按角色和阶段索引的事件,提出受时间约束的关系,并检验这些关系在不同条件下的稳定性;再利用选定关系构建回合级基元变量,并在这些变量上重复分析。每个抽象层级都通过显式测量函数与原始轨迹相连。观察结果和随机化协议实验用于评估所得假设,不同干预实现之间的比较则决定抽象应被保留、细化还是限制。研究给出了可接受归约的有限深度界,识别协议对固定抽象的影响,并刻画实现间的不一致性与抽象误差的组合。有限样本检验使投射干预一致性可操作化,构造示例展示基元构建与抽象细化过程。该表述将此实验方法与语义分类、定性理论归纳及行为模型恢复区分开来,并规定了一套用于发现可泛化行为假设的研究流程;文献层面的新颖性与模型层面的惊奇度分别评估。

关键要点

  1. 01将语言智能体行为科学研究形式化为轨迹抽象层级的学习与检验
  2. 02递归测量角色与阶段事件、时间关系及回合级基元变量
  3. 03通过观察、随机协议实验和干预比较决定保留、细化或限制抽象
  4. 04给出有限深度界并分析协议效应、实现分歧与抽象误差组合
  5. 05区别于语义分类、定性理论归纳和行为模型恢复

解读

尚无解读。

原始英文摘要

arXiv:2610.09237v1 Announce Type: cross Abstract: Scientific studies of language agents need behavioral variables that support hypotheses across tasks and models. We formulate this research problem as learning and testing a hierarchy of trajectory abstractions. A concrete recursive procedure first measures role- and phase-indexed events, proposes temporally constrained relations, and tests their stability across conditions. It then constructs episode-level motif variables from selected relations and repeats the analysis on those variables. Explicit measurement functions connect every abstraction level to the original trajectories. Observations and randomized protocol experiments assess the resulting hypotheses, while comparisons between intervention realizations determine whether an abstraction should be retained, refined, or restricted. We derive a finite-depth bound for accepted reductions, identify protocol effects on fixed abstractions, and characterize realization disagreement and composition of abstraction error. A finite-sample test makes projected intervention consistency operational, and constructed examples illustrate motif construction and abstraction refinement. The formulation distinguishes this experimental approach from semantic taxonomies, qualitative theory induction, and behavior-model recovery. It specifies a proposed research procedure for discovering generalizable behavioral hypotheses, with literature-relative novelty assessed separately from model-relative surprise.

同方向论文 · cs.AI

查看全部 →