端侧语言模型安全性有多脆弱?面向稀疏故障分析的安全关键参数定位
How Fragile Is On-Device Language Model Safety? Localizing Safety-Critical Parameters for Sparse Fault Analysis
Muhammad Zeeshan Karamat · Christiana Chamon Garcia
中文摘要
随着小型语言模型(SLMs)越来越多地部署在资源受限的端侧平台,并作为智能体系统的组成部分,本地存储的模型参数的完整性成为重要的安全问题。研究针对 LLaMA-2-7B-Chat,探讨其安全敏感行为是否集中在参数的稀疏子集内,从而形成可定向分析的简化故障面。提出两种互补的定位方法:低秩安全关联子空间分析,以及参数级安全—效用重要性过滤。两种方法均揭示出网络中存在显著不均匀的安全敏感性,其中 MLP 的 down_proj 一致表现为突出的安全敏感组件,o_proj 贡献较小。基于参数级定位,仅修改 down_proj 中 0.19% 的模型权重即可达到 53% 的 Basic ASR 与 56% 的 GCG ASR,而 tinyBenchmarks 准确率仍保持在 51.6%,未修改基线为 52.2%。上述结果为资源受限、端侧及智能体场景下语言模型的定向故障分析与选择性完整性保护提供了依据。
关键要点
- 01端侧 LLaMA-2-7B-Chat 部署中存在参数完整性的安全隐患,需识别安全敏感子集
- 02提出两种互补定位方法:低秩安全关联子空间分析与参数级安全—效用重要性过滤
- 03MLP 的 down_proj 是最突出的安全敏感组件,o_proj 贡献较小,网络呈现非均匀安全敏感性
- 04修改 down_proj 中仅 0.19% 的权重即触发 53% Basic ASR 与 56% GCG ASR,tinyBenchmarks 准确率仅由 52.2% 降至 51.6%
- 05研究为资源受限、端侧与智能体场景下的定向故障分析与选择性完整性保护提供依据
解读
尚无解读。
原始英文摘要
arXiv:2610.09000v1 Announce Type: cross Abstract: As small language models (SLMs) are increasingly deployed on resource-constrained and on-device platforms, including as components of agentic systems, the integrity of locally stored model parameters becomes an important safety concern. We investigate whether safety-sensitive behavior in LLaMA-2-7B-Chat is concentrated within a sparse subset of parameters, creating a reduced fault surface for targeted analysis. We study two complementary localization methods: low-rank safety-associated subspace analysis and parameter-level safety--utility importance filtering. Both approaches reveal highly non-uniform safety sensitivity across the network, with the MLP down_proj consistently emerging as a prominent safety-sensitive component and o_proj providing a smaller contribution. Using parameter-level localization, modifying only 0.19% of model weights in down_proj yields 53% Basic ASR and 56% GCG ASR, while tinyBenchmarks accuracy remains at 51.6% compared with a 52.2% unmodified baseline. These results motivate targeted fault analysis and selective integrity protection for language models deployed in resource-constrained, on-device, and agentic settings.