跳到正文
原文
Reddit · r/LocalLLaMA· /u/lbgos_Loss783·· 6 天前AI 评分52

用户自建网络安全基准测试:Qwen3.8 27B 等模型的渗透能力实测

I made my own cybersecurity benchmark and ran Qwen3.8 27B, here's how a local model actually does at hacking

AI 导读

作者创建了一个网络安全基准测试,在隔离 Docker 环境中让模型获取 shell 并寻找旗帜,涵盖 pwn、web、crypto、rev、forensics 等 19 个任务,共 544 次评分尝试。

来源:Reddit · r/LocalLLaMA · reddit.com