← All demos

Two models, one experiment

Both ~4-bit, both ~20 GB on disk, both fit the same MacBook — and they're completely different animals.

VS
google/gemma-4-31b-qat

Gemma 4 31B

Dense — one big brain, all of it always on
Every token is computed with all 31 billion parameters.
Total parameters31B
Active per token31B100% of the model
File on disk18.9 GB
QuantizationQAT 4-bittrained with quantization in mind
Maximum brain per token — but the machine reads all ~19 GB of weights for every single token it writes.
qwen/qwen3.6-35b-a3b

Qwen 3.6 35B-A3B

Mixture-of-experts — a big brain that fires in small pieces
The whole 35B lives in memory, but each token only wakes ~3B of it — the "A3B".
Total parameters35B
Active per token3B~9% of the model
File on disk20.4 GB
QuantizationQ4-class 4-bitcompressed after training
Big-model memory cost, small-model speed — each token only reads the ~2 GB of experts it actually uses.

Why one of them feels 10× faster

Generation speed is set by how many weight-bytes the machine must read per token — not by how big the file is.

Gemma 4 31Bdense
31B
Qwen 3.6 35B-A3BMoE
3B
Weights read per token — ~10× less memory traffic for the MoE. Same laptop, same kind of file, wildly different tokens-per-second.

Both models run on a MacBook Pro — M3 Max, 64 GB unified memory — via LM Studio. File sizes as shown in LM Studio. One anecdote, not a benchmark.