Reddit · r/LocalLLaMA· /u/Vasili_Sk·· 2 小时前AI 评分25
为什么 LLM 推理服务每次都重发整段对话历史?上下文能否改成服务端对象?
Why are we still sending the entire conversation to the inference server on every turn? Am I missing something obvious?
AI 导读
提问者指出当前 LLM 推理 API 每轮都要由客户端把整段对话重新发给服务端,迫使 harness 自行实现上下文压缩与 KV cache 匹配,且容易因格式不一致触发 checkpoint 回滚。
来源:Reddit · r/LocalLLaMA · reddit.com