A Qwen3.8-27B user reports 8+ hours of unsupervised agentic work with a 2048-token reasoning budget
The top-voted claim in r/LocalLLaMA today (359 upvotes, 175 comments) is that Qwen3.8-27B is the first local model the poster trusts to run unattended, with the config spelled out in an edit: huihui abliterated Q3_K_XL, reasoning budget capped at 2048 and still coherent at 1024, KV cache at Q8 with 128k context, and a custom chat template plus reasoning-format deepseek to fix tool-call and think-tag generation. Two other threads corroborate the agentic angle, including a run where the model reached a target Wikipedia article in 6 Playwright clicks with no backtracking. This is a single practitioner's report, not a benchmark, but the failure modes named (tool tag generation, compaction) are the ones that actually stop local agent loops.
↳ Follow the thread