时序可解释的可微决策树
Temporally Interpretable Differentiable Decision Trees
Eisuke Hirota · Aarav Sane · Rohan Paleja
中文摘要
可解释性为安全自主决策提供了一种思路,使得智能体的底层决策模型对人类透明。在序贯决策任务中,可微决策树(DDT)是实现可解释性的方法之一,它既能保持策略的可微性,又能为人类提供离散的树形可视化结构。然而,现有 DDT 实现并不适用于序贯决策任务,因为树的单步行为与人类的多步规划之间存在固有不匹配。本工作由此将时间作为可解释性的一个新维度,提出时序可解释性(temporal interpretability)这一概念,并展示了通过动作分块(action chunking)实现的时序抽象如何提升可解释性。具体而言,首先提出两种结合动作分块的策略梯度算法;其次,为保持树的参数高效性,设计了一种基于信息论(information-theoretic)的树结构重构算法,在训练过程中动态调整树结构。在四个仿真环境中的实验表明:以蒸馏得到的动作分块策略来热启动(warm-start)动作分块 DDT,是获得时序可解释树的最有效方式——在四个环境中的三个上,该方法的性能可媲美神经网络策略,同时参数量减少最高达 80%。代码已开源。
关键要点
- 01问题:现有可微决策树(DDT)采用单步决策结构,与人类在序贯决策中的多步规划存在不匹配,缺乏时序层面的可解释性
- 02方法:提出时序可解释性概念,引入动作分块(action chunking)机制,并设计两种新的策略梯度算法以及信息论驱动的树结构重构算法
- 03结果:在四个仿真环境中的三个上,蒸馏热启动的动作分块 DDT 性能可媲美神经网络策略,且参数量减少最高 80%
- 04贡献:开源代码,扩展了 DDT 在序贯决策场景下的可解释性框架,为安全自主系统提供参数高效的可微决策策略
- 05局限:在四个环境中的某一个上未能匹配神经网络策略性能,表明方法并非在所有任务上普遍有效
解读
尚无解读。
原始英文摘要
arXiv:2610.10367v1 Announce Type: cross Abstract: Interpretability offers a solution to safe autonomy by providing transparency into an agent's underlying decision-making model. Within sequential-decision making tasks, differentiable decision trees (DDTs) are one approach to such interpretability, maintaining automatic-differentiable policies while providing humans with a discrete tree-based visualization. Nonetheless, current implementations of DDTs are not well-suited for sequential-decision making domains, as there exists an inherent mismatch between a tree's single-timestep behavior and a human's multi-timestep planning. Our work thus introduces time as a new dimension of interpretability, coined as temporal interpretability, and demonstrates how temporal abstractions via action chunking improve it. We achieve this by first introducing two novel policy gradient algorithms that incorporate action chunking. Additionally, to maintain parameter-efficient trees, we develop an information-theoretic tree restructuring algorithm that modifies the tree during training. Across four simulation environments, we find that warm-starting action chunked DDTs from a distilled action chunked policy is the most effective way to obtain temporally interpretable trees: they match neural network policies in three of the four domains while using up to 80$\%$ fewer parameters. Our code is available at https://github.com/ei5uke/temp-interp.