跳到正文
返回论文列表
cs.CL提交于 已译

InterView-C:VR 化身介导的问卷访谈同步多模态语料库

InterView-C: A Synchronized Multimodal Corpus of VR Avatar-Mediated Survey Interviews

Patrick Schrottenbacher · Leon Hammerla · Lydia Kleine · Doris Stingl · Alexander Mehler

中文摘要

InterView-C 是一个德语多模态语料库,包含 27 场全程在虚拟现实(VR)中进行、由化身代表双方的问卷访谈。语料库将口语交互与同步的行为数据对齐,包括注视、头部与身体运动、面部行为以及手部和手指追踪。其参考转写文本与语言学标注为这种多模态口语交互与以文本为主的 NLP 方法之间提供了可靠接口。该接口至关重要,因为自动语音转写可能扭曲语言学相关信息,而在已转写口语数据上应用的现有资源训练的下游模型还可能面临迁移挑战。InterView-C 因此为全部 54 段录音提供词级对齐、人工后转录的逐字稿、访谈题目时间戳、问卷回答,以及对 1,422 句的否定线索(cue)和否定范围(scope)标注(其中 1,398 句为双重标注;标注员间一致性 α=0.87 用于线索,α=0.81 用于范围)。论文通过实证展示这两类挑战:九个开源 ASR 系统对短封闭回答和数字词的误识别存在系统性偏差;在该语料库转写文本上,现有语料训练的否定模型表现低于使用 InterView-C 标注训练的模型,且性能波动明显。InterView-C 使口语交互的语言学研究得以实现,同时保留其与多模态行为的对齐。

关键要点

  1. 01数据资源:首个在 VR 中由双方化身介导的德语问卷访谈同步多模态语料库,涵盖注视、头/身运动、面部、手指等多模态行为与口语对齐
  2. 02语言学接口:提供 54 段录音的词对齐人工后转录逐字稿、题目时间戳、问卷回答及 1,422 句否定线索/范围双重标注(α 0.87 / 0.81)
  3. 03ASR 实证局限:九个开源 ASR 对短封闭回答和数字词的识别存在系统性偏差,证实自动转写扭曲语言信息
  4. 04否定模型迁移挑战:基于现有语料训练的否定模型在该语料库转写文本上表现低于用 InterView-C 标注训练的模型且不稳定
  5. 05应用价值:支持保留与丰富多模态行为对齐的口语交互语言学研究

解读

尚无解读。

原始英文摘要

arXiv:2610.10145v1 Announce Type: new Abstract: We present InterView-C, a German multimodal corpus of 27 survey interviews conducted entirely in virtual reality, with both interlocutors represented by avatars. The corpus aligns spoken interaction with synchronized behavioral data, including gaze, head and body movement, facial behavior, hand and finger tracking. Its reference transcripts and linguistic annotations provide a reliable interface between this multimodal spoken interaction and predominantly text-based NLP methods. This interface is important because automatically transcribing speech can distort linguistically relevant information, while downstream models trained on existing resources may additionally face transfer challenges when applied to transcribed spoken data. InterView-C therefore provides word-timed and manually post-edited verbatim transcripts for all 54 recordings, interview-item timings, questionnaire responses and negation cue and scope annotations for 1,422 sentences, 1,398 of them doubly annotated ({\alpha}=0.87 for cues; {\alpha}=0.81 for scopes). We demonstrate both challenges empirically: nine open-weight ASR systems disproportionately misrecognize short closed answers and number words, while negation models trained on existing corpora show lower and highly variable performance on our transcribed interviews than a model trained on the InterView-C annotations. InterView-C thus enables linguistic analyses of spoken interaction while retaining their alignment with rich multimodal behavior.

同方向论文 · cs.CL

查看全部 →