用线性注意力 Transformer 执行因果结构学习
Executing Causal Structure Learning with Linear-Attention Transformers
Amartya Roy · Sayar Karmakar
中文摘要
Transformer 能对其输入数据执行算法。论文研究它们是否同样能执行因果发现(casual discovery)。研究聚焦于一种标准连续方法 NOTEARS,它在保持无环性的同时反复更新候选因果图。论文显式构造了一个固定权重的 Transformer,使其前向过程精确复现该方法的一次更新,因而堆叠块可复现其优化轨迹。该 Transformer 在更新之间同时携带当前图与算法的乘子(multiplier)。证明显示,保留乘子对精确执行至关重要,因为不同的乘子值会导致不同的下一步更新。同时给出在固定阶段内,达到目标精度所需的更新次数可预先计算、且舍入误差随深度有界的条件。实验表明:所构造块在浮点精度下与参考更新完全一致;在合成数据及七种已发表基准网络拓扑上,算术回放继承了参考求解器的成功与失败。这将"算法执行精确"与"因果恢复精确"区分开来。相比之下,在给定训练预算下,普通注意力模型既无法可靠地执行更新,也无法迁移到更大的图。梯度训练是否能在该构造所对应的架构类中学会执行器,仍是开放问题。
关键要点
- 01问题:Notebook 通过固定权重 Transformer 能否精确执行 NOTEARS 这类连续因果发现算法的一次更新,且能否与乘子(multiplier)的内部状态保持兼容
- 02方法:显式构造固定权重 Transformer,使其前向过程精确复现 NOTEARS 一次更新,通过堆叠块复现优化轨迹,并证明保留乘子对精确执行的必要性
- 03结果:所构造块与参考更新在浮点精度下完全一致,在合成数据及七个基准网络上的算术回放继承了参考求解器的成败
- 04局限:普通注意力模型在训练预算内无法可靠执行更新或迁移到更大图;梯度训练能否学到该架构类内的执行器仍未解决
解读
尚无解读。
原始英文摘要
arXiv:2610.10395v1 Announce Type: new Abstract: Transformers can execute algorithms on data given in their input. We ask whether they can do the same for causal discovery. We study a standard continuous method that repeatedly updates a candidate causal graph while enforcing acyclicity. We explicitly construct a fixed-weight transformer whose forward pass exactly reproduces one update of this method, so repeated blocks reproduce its optimization trajectory. The transformer carries the current graph and the algorithm's multiplier between updates. We show that retaining the multiplier is essential for exact execution, since different multiplier values can lead to different next updates. We also give conditions under which, within a fixed stage, the number of updates needed to reach a target accuracy can be computed in advance and rounding errors stay bounded as depth grows. Experiments show that the constructed block agrees with a reference update to floating-point precision, while arithmetic replay on synthetic data and seven published benchmark network topologies inherits the reference solver's successes and failures. This separates accurate algorithm execution from accurate causal recovery. In contrast, the ordinary attention models tested under our training budgets do not reliably execute the update or transfer to larger graphs. Whether gradient training can learn an executor in the architecture class of the construction remains open.