跳到正文
原文
Reddit · r/LocalLLaMA· /u/exaknight21·· 5 小时前AI 评分12

8 卡 T4 服务器部署 Qwen3.5-9B 推理速度求助:vLLM 还是 llama.cpp 更快?

I have an ESC4000 G3 with 8x T4s in it - what is the fastest way I can deploy Qwen3.5-9B for about 10-15 users concurrently: currently using llama.cpp

AI 导读

网友 exaknight21 在 Asus ESC4000 G3(128 GB DDR4、8 张 T4)上尝试部署 Qwen3.5-9B-AWQ-INT8 服务 10–15 人并发:当前用 llama.cpp 开 8 个实例跑 Q5_K_XL 量化版本,提供 44 个 slot,单 slot 持续生成速度 25–40 tps,prefill 速度 1400 tps。

来源:Reddit · r/LocalLLaMA · reddit.com