跳到正文
原文
Reddit · r/LocalLLaMA· /u/Vasili_Sk·· 2 小时前AI 评分25

为什么 LLM 推理服务每次都重发整段对话历史?上下文能否改成服务端对象?

Why are we still sending the entire conversation to the inference server on every turn? Am I missing something obvious?

AI 导读

提问者指出当前 LLM 推理 API 每轮都要由客户端把整段对话重新发给服务端,迫使 harness 自行实现上下文压缩与 KV cache 匹配,且容易因格式不一致触发 checkpoint 回滚。

来源:Reddit · r/LocalLLaMA · reddit.com