跳到正文
原文
Reddit · r/LocalLLaMA· /u/turtleninja99·· 3 天前AI 评分18

在 64 GB Mac mini 上做 MoE SSD 流式加载:Qwen Flash Next 解码时 GPU 仍有 27% 时间在等专家层

MoE SSD streaming on a 64 GB Mac mini: GPU still waits 27% of decode on experts. Ideas?

AI 导读

用户在 64 GB Mac mini(M5)上运行 Qwen Flash Next 的 q4 量化版本,由于模型放不进内存,采用将热门专家留在缓存、其他专家从 SSD 流式加载的方案,测得 17.5 t/s 解码和 390 t/s prompt 处理速度,热缓存命中率为 75%。

来源:Reddit · r/LocalLLaMA · reddit.com