← All demos

One name, many brains

“Gemma 4” and “Qwen 3.6” aren't single models — they're families. Same training recipe, different brain sizes. Bigger brain → smarter, but it needs more memory. Circle area = parameter count.

Total parameters (what memory must hold) Active / effective per token (what sets the speed)

“Needs” = model file + a 32k-token context + runtime overhead. Dense models use every parameter for every token; MoE models keep the whole brain in memory but only fire a few experts per token — big-model memory cost, small-model speed. Gemma's E-series (PLE) plays a similar trick with per-layer embeddings: the file holds the full count, but far fewer parameters do the work. Specs from the official Gemma 4 and Qwen 3.6 model cards (April 2026 releases).