Reddit · r/LocalLLaMA· /u/theexile1337·· 5 天前AI 评分52
ninfer vs llama.cpp 在 RTX 5090 上运行 Qwen3 27B 的性能与质量对比评测
2.3x faster Qwen3.8 27B on a 5090: ninfer vs llama.cpp, 4 setups, same prompt - speed and quality tested
AI 导读
在 RTX 5090 上测试 Qwen3 27B 运行 ninfer 和 llama.cpp 的性能,结果显示 ninfer 比 llama.cpp 快约 2.3 倍,但 llama.cpp 开启 MTP 后速度从 68 t/s 提升至 141 t/s,接近 ninfer 水平。
来源:Reddit · r/LocalLLaMA · reddit.com