跳到正文
返回论文列表
cs.CV提交于 已译

通过视觉-语言模型响应特征检测对抗图像

Detecting Adversarial Images through Response Profiles of Vision-Language Models

Arash Vashagh · Roozbeh Razavi-Far

中文摘要

对抗扰动可以改变冻结状态的视觉-语言模型(VLM)的预测结果,同时使其置信度和图像-文本相似性模式表面上看起来仍然合理。本研究探讨能否基于图像与一组通用语义提示之间更广泛的交互方式来识别对抗输入。检测器利用类别级统计量、提示之间的关系、相对干净参考分布的偏差,以及在弱图像变换下的稳定性来汇总这些响应,生成一个紧凑的响应特征,并由一个轻量级模型进行分类,而 VLM 保持固定不变。该方法在多个公开图像数据集、多种 CLIP 风格的视觉骨干网络,以及涵盖基于梯度的、基于优化的、自动化的和空间性的多种攻击上进行了评估。检测器在针对特定攻击的设定下表现出强大的判别能力,并在面对训练中未见的攻击时仍保持显著性能。在受控的检测器专属协议下,响应特征表示优于所评估的嵌入几何基线。进一步分析表明,各特征组提供互补信息,且该方法在提示配置变化下仍然有效。同时考察了推理成本以及针对检测器感知型自适应攻击的鲁棒性。总体而言,结果表明跨语义提示的响应模式为冻结 VLM 中的对抗图像检测提供了一种有用的互补信号。

关键要点

  1. 01问题:对抗扰动能改变冻结 VLM 的预测,但置信度与图像-文本相似性仍看似合理,难以直接察觉。
  2. 02方法:构建由类别统计、提示关系、相对干净参考的偏差、弱变换稳定性组成的响应特征,用轻量分类器检测,VLM 保持冻结。
  3. 03结果:在多种数据集、CLIP 骨干、跨多种攻击(含未见攻击)下判别能力强,特征组互补,优于嵌入几何基线。
  4. 04局限:检测器面临检测器感知型自适应攻击的挑战,并受推理成本约束。

解读

尚无解读。

原始英文摘要

arXiv:2610.10436v1 Announce Type: new Abstract: Adversarial perturbations can alter the predictions of frozen vision-language models (VLMs) while leaving their confidence and image--text similarity patterns seemingly plausible. We investigate whether we can identify adversarial inputs based on the broader way an image interacts with a collection of general semantic prompts. Our detector summarizes these responses using category-level statistics, relationships among prompts, deviations from clean reference distributions, and stability under weak image transformations, producing a compact response profile that is classified by a lightweight model while the VLM remains fixed. We evaluate the approach on multiple public image datasets, several CLIP-style visual backbones, and a range of gradient-based, optimization-based, automated, and spatial attacks. The detector achieves strong discrimination in attack-specific settings and retains substantial performance when evaluated on attacks not seen during training. Under a controlled detector-specific protocol, the response-profile representation outperforms the evaluated embedding-geometry baselines. Additional analyses show that the feature groups provide complementary information and that the method remains effective under variations in the prompt configuration. We also examine inference cost and performance against detector-aware adaptive attacks. Overall, the results indicate that response patterns across semantic prompts provide a useful complementary signal for adversarial image detection in frozen VLMs.

同方向论文 · cs.CV

查看全部 →