Derivative of Qwen/Qwen3.6-35B-A3B, compressed to a hard 13.5 GiB byte budget so that the full model plus 64k context fits in 16 GiB of GPU memory (measured on unified memory; see the fit section for what that does and does not prove). If you have more memory, use the standard tiers instead: MagicQuant hybrids (Q4/Q5/Q6) or the ROCmFPX build.

The experiment:

The smallest standard tier of this model is 20.2 GiB (21.7 GB). It cannot fit in 16 GiB at any context length, so this repo asks a different question: what is the best 35B-A3B you can have if the budget is fixed at 16 GiB total, including 64k of context? (Sizes on this card are GiB, 1024-based, throughout.)

The context arithmetic is what makes it plausible at all. Qwen3.6-35B-A3B is a hybrid-attention model: only 10 of its 40 layers are full attention (2 KV heads x 256 head dim); the other 30 are linear-attention layers whose state does not grow with context. KV cache at 64k is therefore about 1.25 GiB at f16, several times smaller than a dense model of this size. 16 GB minus KV minus runtime buffers leaves roughly 13.5 GiB for weights, a 0.20 size ratio versus BF16, well below the Q4 band.

The weights were fitted to that budget with MagicQuant v2’s size-target search: an exact per-tensor knapsack (753 tensor assignments, imatrix-calibrated) under a hard byte ceiling, verified against real perplexity rather than a proxy. The file landed at 13.51 GiB against a 13.5 GiB request.

A second copy of the model then went through quantization-aware training (QAT): frozen-mode LoRA trained against this exact quantization layout, merged, and re-packed at the identical 753-tensor allocation. Both copies are published because they win on different workloads (measurements below).