Reddit · r/LocalLLaMA· /u/Training_Visual6159·· 5 天前AI 评分37
Strata 发布 MoE 缓存方案,性能超越 llama.cpp 5-10 倍
Imma just say it, Strata absolutely clowned llama.cpp
AI 导读
Strata 用约两周时间实现了 MoE 缓存优化,带来 5-10x prefill 和 3-4x decode 性能提升。方案基于 Qwen-3.8-flash-next 模型,在 8-16gb 显卡 + 64gb 内存配置下达 2000/70 t/s+ 速度,比 27b 模型更快。
来源:Reddit · r/LocalLLaMA · reddit.com