跳到正文
返回论文列表
cs.CL提交于 已译

Itgan 在 NADI 2026 共享任务中的方案:面向鲁棒、混合方言与代码转换阿拉伯语 ASR 的 Whisper 参数高效适配

Itgan at NADI 2026 shared task: Parameter-Efficient Whisper Adaptation for Robust, Mixed-Dialect and Code-Switched Arabic ASR

Ibrahim Almajai

中文摘要

介绍 Itgan 团队参加 NADI 2026 三个 ASR 子任务的系统,分别为鲁棒国家级 ASR(1.1)、混合方言 ASR(1.2)以及突尼斯语代码转换 ASR(1.3)。三者共用同一套配方——在消费级 GPU 上用 LoRA 适配 Whisper,各子任务在此基础上各加一项改动。子任务 1.1 在测试时已知方言标签,从一个合并适配器继续微调得到的每方言专用模型带来最大提升,提交系统达到 57.1% 的国家级平均 WER。评测后对冻结编码器特征做线性探针,可在无标签时路由语句,恢复 oracle 路由 44% 的收益。子任务 1.2 中,基模型的选择比适配器容量更关键;只有加入一个去相关成员后,系统融合才有帮助,最终 WER 为 46.7%。子任务 1.3 系统以 14.49% WER 排名第二,在领先提交中 CER 最低,为 5.38%。其中最后 0.60 个 WER 点的提升无需再训练,主要来自若干独立训练运行在权重空间上的精确平均,剩余部分由 ROVER 投票带来。所有比较均使用配对自助检验,并报告了八个未奏效的方向。

关键要点

  1. 01针对 NADI 2026 三个阿拉伯语 ASR 子任务,统一采用 Whisper + LoRA 在消费级 GPU 上适配,各子任务叠加不同改进
  2. 02国家级 ASR(1.1)中方言标签已知时,从合并适配器继续微调的每方言专家模型收益最大,提交系统达到 57.1% 平均 WER
  3. 03混合方言 ASR(1.2)中基模型选择比适配器容量更关键,系统融合需加入去相关成员才有帮助,WER 为 46.7%
  4. 04突尼斯语代码转换 ASR(1.3)以 14.49% WER 排名第二,CER 5.38% 为领先提交中最低,最后 0.60 WER 点来自权重空间精确平均与 ROVER 投票
  5. 05所有比较均做配对自助检验,并明确列出八个未奏效的方向,体现负面结果的报告

解读

尚无解读。

原始英文摘要

arXiv:2610.09934v1 Announce Type: new Abstract: We describe the Itgan systems for the three ASR subtasks of NADI 2026, namely robust country-level ASR (1.1), mixed-dialect ASR (1.2), and Tunisian code-switched ASR (1.3). All three share one recipe, Whisper adapted with LoRA on consumer GPUs, and each was carried by a different addition to it. On 1.1, where the dialect label is given at test time, per-dialect specialists continued from a pooled adapter gave the largest gain, and the submitted system reached 57.1% country-average WER. A post-evaluation linear probe on frozen encoder features routes utterances without the label and recovers 44% of what oracle routing gives. On 1.2 the choice of base model mattered more than adapter capacity, and system combination helped only once we added a decorrelated member, reaching 46.7% WER. On 1.3 our system placed second at 14.49% WER with the lowest CER among the leading submissions, 5.38%. Its last 0.60 WER points came without further training, mostly from an exact weight-space average of independently trained runs, with ROVER voting adding the remainder. Every comparison carries a paired-bootstrap test, and we report eight directions that did not work.

同方向论文 · cs.CL

查看全部 →