跳到正文
原文
Reddit · r/LocalLLaMA· /u/evp-cloud·· 4 天前AI 评分75

Qwen3.8 27B 在单张 AMD R9700 上的性能:262K 上下文、50 万 tokens 可重用缓存

Qwen3.8 27B | 1 x R9700: 262K context, half a million tokens of reusable cache, ~180 tok/s. And yes, let's talk about the "3-bit" :)

AI 导读

作者发布了 Qwen3.8 27B 模型在单张 AMD Radeon AI PRO R9700(32GB)上的运行结果,使用 3-bit 权重(仅大投影矩阵)和 speculative decoding,可重用缓存达到 569,878 tokens,每个请求支持 262,144 tokens 上下文。

来源:Reddit · r/LocalLLaMA · reddit.com