跳到正文
返回论文列表
cs.CL提交于 已译

Prompt 应该做更多:检索指令对嵌入模型(embedding models)的影响

Your Prompt Should Do More: Effects of Retrieval Instructions in Embedding Models

Amanda Myntti · Jenna Kanerva · Veronika Laippala · Filip Ginter

中文摘要

基于 Prompt 的嵌入模型(embedding models)近期在检索任务中受到关注,检索时会在 prompt 中提供详细的检索指令。多个新数据集与研究表明,当前嵌入模型在可靠遵循这些指令方面仍存在困难。本研究分析在非对称检索(asymmetric retrieval)任务中,指令究竟如何影响检索查询(query)的表示。实验表明,即便在简单任务指令下,当评估中引入查询侧干扰项(query-side distractors)时,模型也会失效。本文假设该现象源于当前嵌入模型的训练设定与评估方式,并证明在微调中加入查询侧干扰项可带来显著提升,同时对其他任务的影响极小。

关键要点

  1. 01问题:现有嵌入模型在非对称检索中难以可靠遵循详细的检索指令
  2. 02发现:即便任务指令简单,加入查询侧干扰项后模型也会失效
  3. 03归因:该失败行为源于当前嵌入模型的训练与评估设定
  4. 04方法:在微调中加入查询侧干扰项进行训练
  5. 05结果:该方法带来实质性提升,且对其他任务影响极小

解读

尚无解读。

原始英文摘要

arXiv:2610.10508v1 Announce Type: new Abstract: Prompted embedding models have recently received increasing attention, particularly for retrieval, where detailed retrieval instructions are provided as part of the retrieval prompt. Several new datasets and studies have examined this setting, showing that the current embedding models often struggle to follow such instructions reliably. In this paper, we study the mechanism of how instructions actually affect the representations of retrieval queries in asymmetric retrieval tasks. We show that models can fail to follow even simple task instructions when query-side distractors are included in the evaluation. We hypothesize that this behavior is driven by the training setup of current embedding models and their evaluation, and show that fine-tuning with added query-side distractors leads to substantial improvements, with minimal effect on other tasks.

同方向论文 · cs.CL

查看全部 →