基于标量伴随匹配的 Q 学习
Q-Learning with Scalar Adjoint Matching
Yonghoon Dong · Minsung Yoon · Jaehyuk Kim · Jungwoo Park · Changyeon Kim · Jinwoo Shin
中文摘要
流策略(flow policy)能够捕捉丰富多样的动作分布,近年来用离线强化学习(off-policy RL)对其进行微调以超越示范水平逐渐成为研究热点。然而,基于学习到的价值函数对流策略进行微调并非易事,因为策略需要经过多步流生成才能产出最终动作。伴随匹配(adjoint matching)提供了一种原则性的方法,通过将最终动作处的价值信息反向传播至每一步流,从而直接更新流模型本身;但该方法在每一步都需要对策略做一次向量-雅可比积(vector–Jacobian product),其计算成本随流步数和策略规模线性增长。观察到预训练流策略的批量平均速度雅可比矩阵集中于其对角线,由此推导出一个闭式的标量伴随(scalar adjoint),它仅将最终动作处的价值梯度按流时间进行缩放,从而消除了逐步向量-雅可比积的开销。同时发现在标量伴随下,控制评论器(critic)在策略生成动作处的价值尤为重要。基于上述发现,提出带标量伴随匹配的 Q 学习(SQAM),将标量伴随与对这些动作的价值惩罚相结合。SQAM 的性能提升集中在 OGBench 四个最难的任务上,在各任务上成功率均超过最强基线 18 至 35 个百分点。为验证 SQAM 是否可扩展至大规模预训练策略,还在真实双臂机器人上对一个视觉-语言-动作(vision-language-action)策略进行了微调,在全部三项任务上均优于监督微调。
关键要点
- 01流策略用 RL 微调的核心瓶颈:每步向量–Jacobian 积,代价随流步数和策略规模线性增长
- 02发现预训练流策略的批量平均速度雅可比矩阵集中于对角线,据此推导闭式标量伴随,消除逐步向量–Jacobian 积
- 03标量伴随下需要专门控制 critic 在策略生成动作上的价值,提出价值惩罚与之结合形成 SQAM
- 04SQAM 在 OGBench 四个最难任务上成功率超过最强基线 18 至 35 个百分点
- 05在真实双臂机器人 VLA 策略的三项任务上均优于监督微调
解读
尚无解读。
原始英文摘要
arXiv:2610.10437v1 Announce Type: cross Abstract: Flow policies capture rich and diverse action distributions, and fine-tuning them with off-policy RL to improve beyond the demonstrations has drawn growing interest. However, fine-tuning a flow policy against a learned value function is not trivial, because the policy generates its action over many flow steps. Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size. We observe that the batch-averaged velocity Jacobian of pretrained flow policies concentrates on its diagonal. Motivated by this finding, we derive a closed-form scalar adjoint that scales the value gradient at the final action by the flow time, eliminating the per-step vector--Jacobian products. We further find that controlling the critic's value at policy-generated actions is particularly important under the scalar adjoint. Based on these findings, we propose Q-learning with Scalar Adjoint Matching (SQAM), which combines the scalar adjoint with a value penalty at those actions. SQAM's gains concentrate on the four hardest OGBench domains, where its success rate exceeds that of the strongest baseline in each domain by 18 to 35 percentage points. To test whether SQAM extends to large pretrained policies, we also fine-tune a vision-language-action policy on a real bimanual robot. SQAM improves over supervised fine-tuning on all three tasks.