跳到正文
返回论文列表
cs.AI提交于 已译

Route-Verify-Vote:面向混合领域推理的程序条件自一致性方法

Route-Verify-Vote: Procedure-Conditioned Self-Consistency for Mixed-Domain Reasoning

Xinchen Xiao

中文摘要

组合泛化仍是语言模型面临的难题——当模型需要以不熟悉的方式组合已掌握的推理操作时尤为如此。SCoRE(Scenario-Based Commonsense Reasoning Evaluation)2026 在三个训练中未出现的混合领域上测试这一能力,要求模型为每道题给出完整的正确选项集合。论文提出 Route-Verify-Vote (RVV) 框架,在不更新模型参数的前提下实现程序条件的自一致性。Route 利用题目提供的领域标签选择推理程序,引导模型表示并应用相应约束;Verify 提示模型依据这些约束评估每个选项;Vote 聚合完整答案集,并将额外采样分配给最高票两个集合票数差较小的题目。每道题的所有样本遵循同一领域特定程序。在官方测试集上,对每道题采样 16 个答案集进行投票,精确集合准确率达 74.6%;自适应 RVV 达到 77.3%;在选定领域路线上融合多模型,准确率提升至 79.4%。该系统在参赛系统中最终排名第二。结果表明,领域特定推理程序与答案集分歧可作为混合领域推理中分配推理时计算的有用工具。

关键要点

  1. 01问题:语言模型在训练未出现的混合领域上难以将熟悉的推理操作以新方式组合,SCoRE 2026 任务要求给出完整正确选项集
  2. 02方法:Route-Verify-Vote 框架,Route 按领域标签选程序,Verify 逐项校验约束,Vote 聚合完整答案集并对票差小的题追加采样
  3. 03结果:16 样本投票达 74.6%,自适应 RVV 达 77.3%,多模型融合特定领域路线后达 79.4%,最终排名第二
  4. 04关键机制:所有样本共享同一领域特定程序保证一致性,答案集分歧度驱动推理时算力分配

解读

尚无解读。

原始英文摘要

arXiv:2610.08814v1 Announce Type: cross Abstract: Compositional generalization remains challenging when language models must combine familiar reasoning operations in unfamiliar ways. The Scenario-Based Commonsense Reasoning Evaluation (SCoRE) 2026 tests this ability on three mixed domains absent from training and requires models to identify the complete set of correct options for each question. We introduce Route-Verify-Vote (RVV), a framework for procedure-conditioned self-consistency that uses language models without parameter updates. Route uses the provided domain label to select a reasoning procedure that guides the model in representing and applying the relevant constraints. Verify prompts the model to assess each option against those constraints. Vote aggregates complete answer sets and allocates additional samples to questions with a small vote-count margin between the two most frequent sets. Samples for each question follow the same domain-specific procedure. On the official test set, voting over 16 sampled answer sets per question achieves an exact-set accuracy of 74.6%. Adaptive RVV reaches 77.3%, and combining models on selected domain routes raises accuracy to 79.4%. The final system ranked second among participating systems. These results support domain-specific reasoning procedures and answer-set disagreement as useful tools for allocating inference-time computation in mixed-domain reasoning.

同方向论文 · cs.AI

查看全部 →