Reddit · r/LocalLLaMA· /u/BringTea_666·· 3 天前AI 评分25
单卡 RTX5090 跑本地 LLM 推理遇瓶颈:解码速度足够快,CPU 反而成为智能体编码的真正瓶颈
Practical limit hit. Decoding so fast that tool calls (cpu) starting to become real limit not decode or prefill. Single RTX5090. Porting Kenshi to Godot project.
AI 导读
开发者将 Kenshi 项目移植到 Godot,使用单张 RTX5090 跑本地 LLM 推理服务,GPU 平均解码约 700t/s 却无法再提速;启动 25 个智能体并发占满 12 个服务槽位后,任务卡在 tool call 阶段,CPU 满载。
来源:Reddit · r/LocalLLaMA · reddit.com