跳到正文
返回论文列表
cs.CL提交于 已译

逻辑与人设解耦:边缘大语言模型 Agent 对上下文污染的结构性免疫

Decoupling Logic from Persona: Structural Immunity of Edge LLM Agents to Context Pollution

Masaaki Nakatsu (AO · Inc. / OrbLabs AG) · Reno Wang (AO · Inc.)

中文摘要

边缘设备上的小语言模型 Agent 需要在同一上下文窗口内同时承载人设(persona)与逻辑推理,而该窗口会被对话历史与人设指令填满。研究关注当历史冗长、带误导性且人设内容密集时(即人设—逻辑干扰),Agent 逻辑部分的表现,并提出解耦架构 AO-DA:基于同一 INT4 量化基模型,配合可热插拔的 LoRA 适配器,将逻辑推理("What")与人设表达("How")分离到两条推理路径。逻辑路径仅接收当前轮核心输入,输出可验证的结构化状态(Micro-State);人设路径在完整历史下将其按角色风格渲染。在 Apple M2 笔记本上对同基模型进行消融实验(Llama-3.1-8B-Instruct 与 Gemma-3-4B-it,4-bit;480 次运行,涵盖 4 个污染等级 × 3 个分支 × 2 个任务 × 2 个人设 × 5 个种子),结果包括:(i)解耦逻辑路径对污染结构性免疫:其 prompt 稳定在 180(Llama)或 167(Gemma)个 token,而混合单次路径的 prompt 从 242 增至 1,203,且输出在各级污染下完全字节一致(40/40);(ii)混合单次路径性能随污染单调下降,Llama 上复合逻辑得分由 0.669 降至 0.150,Gemma 上由 0.487 降至 0.150,主要原因是无法输出要求的结构化格式(Llama 在最高两级下占 80–95%,Gemma 占 100%);(iii)在将相同污染输入解耦逻辑路径后,专适配器、专格式路径在 8B 模型上仍优于单次路径(失败率 0–20% 对比 80–95%;配对 Δ +0.30 至 +0.50,Cliff's δ 0.50–0.85,Holm 校正后 p ≤ 0.03),但在 4B 模型上两者均崩溃。分离的代价仅在话题首轮多一次解码(Llama 上 28.2 s 对比 18.2 s),收益则是 1.7 ms 内的热插拔人设切换且无需重跑逻辑路径。代码、评分细则、测试样本、适配器与日志已开源。

关键要点

  1. 01问题:边缘小模型 Agent 在同一上下文窗口内同时承担人设与逻辑推理,长且污染的历史会引发人设—逻辑干扰,损害逻辑能力
  2. 02方法:AO-DA 解耦架构在同一 INT4 基模型上用可热插拔 LoRA 适配器分两条路径——逻辑路径只接收本轮输入输出 Micro-State,人设路径负责风格渲染
  3. 03结果:解耦逻辑路径 prompt 在 Llama 上稳定在 180 token、Gemma 上 167 token,输出在四级污染下 40/40 完全字节一致,而混合单次路径性能随污染单调下降至 0.150
  4. 04局限:在 4B 模型上解耦与单次路径在高污染下均崩溃,分离在话题首轮多耗约 10 s 解码时间
  5. 05效果:8B 模型上失败率从 80–95% 降至 0–20%,并支持 1.7 ms 人设热插拔

解读

尚无解读。

原始英文摘要

arXiv:2610.09772v1 Announce Type: new Abstract: Small language-model agents on edge devices must hold a persona and reason correctly at once, inside one context window that fills with conversational history and persona instructions. We study what happens to the logical part of such an agent when that history is long, misleading and persona-heavy (persona-logic interference), and present a Decoupling Architecture (AO-DA) that separates logical inference ("What") from persona expression ("How") into two inference paths on one INT4 base model with hot-swappable LoRA adapters. The logic path receives only the core turn and emits a verifiable structured state (Micro-State); the persona path renders it in character with the full history. In same-base-model ablations on an Apple M2 laptop (Llama-3.1-8B-Instruct and Gemma-3-4B-it, 4-bit; 480 runs over 4 pollution levels x 3 arms x 2 tasks x 2 personas x 5 seeds) we find: (i) the decoupled logic path is structurally invariant to pollution: its prompt stays at 180 (Llama) or 167 (Gemma) tokens while the mixed single-pass prompt grows from 242 to 1,203, and its outputs are byte-identical across levels (40/40); (ii) the mixed single pass degrades monotonically (composite logic score 0.669 to 0.150 on Llama, 0.487 to 0.150 on Gemma), mostly by failing to emit the required structured output (80-95% of runs on Llama, 100% on Gemma at the two highest levels); (iii) with the same pollution fed into the decoupled logic path, the dedicated-adapter, dedicated-format path is still more robust than the single pass on the 8B model (failure 0-20% vs 80-95%; paired $\Delta$ +0.30 to +0.50, Cliff's $\delta$ 0.50-0.85, Holm-adjusted $p \le 0.03$) but not on the 4B model, where both collapse. Separation costs one extra decode on a topic's first turn (28.2 s vs 18.2 s on Llama) and buys persona hot-swapping in 1.7 ms without re-running the logic path. Code, rubric, fixtures, adapters and logs are released.

同方向论文 · cs.CL

查看全部 →