Hi all, if you’ve not heard, there’s been a Caveman compression version of Qwen 3.6-27b and Qwen 3.6-35b-3ab.
https://huggingface.co/ProCreations/grug-27b-gguf
https://huggingface.co/ProCreations/grug-35b-v2
https://huggingface.co/ProCreations/grug-35b-v2-gguf
I’ve been playing around with 35B and I am able to run it on my ancient Quadro P1000 4gb at ~10tok/s.
But beyond that, the quality of the reasoning and the token discipline / output is actually higher than a stock, in my opinion.
You can read the benchmarks above; I am also uploading a HTML file here for your consideration.
(Sorry for the Limewire link; I dunno where else to share throw-away files. It’s HTML)
Anyway…I’m doing more testing right now…but so far, this is a good cook.
Or -
Grug good. Me like.


Oh and if anyone is interested in my local settings
-t 12 ^ -tb 12 ^ -ngl 28 ^ --n-cpu-moe 41 ^ --no-mmap ^ --mlock ^ --cache-type-k q4_0 ^ --cache-type-v q4_0 ^ -c 16384 ^ -b 256 ^ -ub 128 ^ --host 0.0.0.0 ^ --port %PORT% ^ --ui-mcp-proxyTurbo quant is not supported on this GPU (too old),
Slight revision: on my rig (i7-8700, 32gb, Quadro p1000 4gb ddr5), this gives a nice 12 tok/s. Unfortunately, it’s still hostile to my 1L box and very quickly thermally swamps the CPU (while gpu sits at 56 degrees, that little shit). C’est la vie.
-t 8 ^ -tb 8 ^ -ngl 99 ^ --n-cpu-moe 38 ^ --flash-attn on ^ --no-mmap ^ --mlock ^ --cache-type-k q4_0 ^ --cache-type-v q4_0 ^ -c 16384 ^ -b 256 ^ -ub 128 ^ --host 0.0.0.0 ^ --port %PORT% ^ --ui-mcp-proxyBe aware, q4_0 KV quantization really borks models.
But the K is far more sensitive than the V. Try q5_1 for the K while leaving the V at q4_0 or q4_1; vram usage will be almost the same, but it should work dramatically better.
As for throttling, try disabling turbo on the 8700.
It doesn’t actually need turbo clocks for these models. The t/s loss I get from doing that on my rig is very modest.
And like others suggested, try the QAT release. You might try the ik_llama.cpp for while you’re at it, at it should be faster with MoEs like this.