输入盲控制在多选题评估中为层程序带来显著 oracle 上限空间
Input-Blind Controls Produce Substantial Oracle Headroom for Layer Programs in Multiple-Choice Evaluation
Yibei Guo · Rui Liu
中文摘要
自适应计算旨在通过为每个输入定制执行过程来改进语言模型推理。针对层程序,oracle 评估在实用选择器出现前利用已知答案估算这种灵活性的潜在收益。但仅靠选择带来的收益并不能解释所选程序为何有效。本研究区分二者:基于两个模型上的 32 个层跳过与重复程序及 4,413 道多选题,比较它们相对于不使用评估提示选定的固定动作的收益,以及在相同位置上输入盲扰动的收益,并在另一提示上重新评估选择。在选项顺序共享的情况下,这些控制在 Qwen3-4B-Base 和 Llama-3.1-8B 上分别产生 10.2–11.8 和 15.6–19.4 个百分点的上限空间,在每个模型三次随机方向抽样中均超过真实程序的 9.0 和 10.1。它们仅匹配答案变化率,且排序取决于菜单:事后比较中,真实程序在 Llama 仅含重复的菜单中每次抽样均领先。一个更小的 KL 校准比较(含输入相关控制)在点位估计上偏向真实程序,但校正后检验无定论。固定字母偏移产生相近量级的上限空间。旋转选项显著降低两类程序的上限空间,但保留 1.4–2.3 与 3.7–4.5 的正向真实程序与控制性差值;其大小与统计支撑取决于进一步调整和参考。补充的生成式回答测试发现,搜索选定的程序在改写后对为其他问题选定的程序保持 26.0 点的优势,但缺少对照。这些结果表明,在共享选项顺序的提示下,显著上限空间可以持续存在,但并不确立特定于所选层计算的收益;针对这些控制的排序均无法识别该收益。
关键要点
- 01问题:oracle 收益评估能估计自适应推理的潜在增益,但无法解释所选层程序为何有效
- 02方法:在 Qwen3-4B-Base 和 Llama-3.1-8B 上,对 32 个层跳过/重复程序比较其真实收益与输入盲控制在相同位置、另一提示上的收益
- 03结果:输入盲控制在共享选项顺序下产生 10.2–19.4 点的上限空间,超过真实程序的 9.0 与 10.1 点;旋转选项后差值缩至 1.4–4.5 点
- 04局限:小规模 KL 校准比较校正后检验无定论,生成式回答测试缺少安慰剂对照,难以确立所选层计算特有的收益
解读
尚无解读。
原始英文摘要
arXiv:2610.10368v1 Announce Type: new Abstract: Adaptive computation aims to improve language-model inference by tailoring execution to each input. For layer programs, oracle evaluations use known answers to estimate the potential gain from this flexibility, before a practical selector is available. However, a gain from selection does not by itself explain why the chosen programs help. This study examines this distinction using 32 layer-skipping and repetition programs on two models and 4,413 multiple-choice items. The analysis compares their gains over a fixed action selected without the evaluation prompt with those of input-blind perturbations at the same sites, re-evaluating selections on another prompt. With shared option order, the controls give 10.2-11.8 and 15.6-19.4 percentage points of headroom on Qwen3-4B-Base and Llama-3.1-8B, exceeding the real programs' 9.0 and 10.1 in all three random-direction draws per model. They match answer-change rate only, and the ordering depends on the menu: in post hoc comparisons, real programs lead on Llama's repeat-only menu in every draw. A smaller KL-calibrated comparison, including an input-dependent control, favours real programs in point estimate, with inconclusive corrected tests. Fixed letter offsets produce headroom of similar scale. Rotating options sharply reduces both families' headroom, while leaving positive real-minus-control differences of 1.4-2.3 and 3.7-4.5 points; their magnitudes and statistical support depend on further adjustments and the reference. A supplementary generated-answer test finds that search-selected programs keep a 26.0-point advantage over programs selected for other problems after rewording, without a placebo comparison. These results show that substantial headroom can persist across prompts with shared option order without establishing a benefit specific to the selected layer computation; neither ordering against these controls identifies that benefit.