Reddit · r/LocalLLaMA· /u/Competitive-Scar-627·· 6 天前AI 评分20
本地大模型推理:6GB GPU + 24GB RAM 如何提升运行速度
Model weight inferencing
AI 导读
用户在 RTX 4050 6GB GPU + 24GB RAM 环境下运行 Qwen2.5 27B 模型时速度过慢,咨询 weight inferencing 是否有助于提升推理速度,并寻求具体尝试方法。
来源:Reddit · r/LocalLLaMA · reddit.com