跳到正文
原文
Reddit · r/LocalLLaMA· /u/W61k3r·· 6 天前AI 评分73

用户分享 Qwen3.8-27b 量化模型:24GB 显卡实测 36-41 tokens/s

Tuned/abliterated Qwen3.8-27b into a 24gb card 262k guff using the newest unreleased version of LexiPanel. It's fast with reliable draft acceptance. Made for 7900xtx but should work on whatever 24gb card with this setup and headless. Doesn't get dumber while coding like most of the other fine-tunes.

AI 导读

用户使用 LexiPanel 将 Qwen3.8-27b 量化并 abliterate 到 24GB 显卡,支持约 262k 上下文和 MTP draft。在 RX 7900 XTX 上实测 110k-170k tokens 上下文时解码速度 36-41 tokens/s,MTP draft 接受率中位数 85%。

来源:Reddit · r/LocalLLaMA · reddit.com