← All demos

Can I run it?

Pick hardware and a model — everything below is computed from two numbers: bytes of weights and memory bandwidth.

Quantization = storing each weight in fewer bits. Fewer bits → smaller file + faster output, slightly worse answers. Names like Q4_K_M are llama.cpp presets: Q4 ≈ 4-bit.
Parameters ("weights") = the model's learned numbers. More → more capable, but bigger and slower. 14B = 14 billion.
MoE models only run a few "expert" chunks per token — memory cost of the full model, speed of a much smaller one.
Context = how much text the model can see at once (a token ≈ ¾ of a word). Its working memory of that text — the KV cache — grows linearly with it.
Weights KV cache Runtime overhead Free fast memory
Hover any legend term for what it means. Everything must fit left of the capacity tick to run at full speed.
File on disk
—
params × bytes/param
Generation speed
—
Prompt processing
—
Verdict
—

The one formula doing all the work: tokens/sec ≈ bandwidth ÷ active-weight-bytes × efficiency. Estimates assume llama.cpp-style inference; GPU offload spills to DDR5 system RAM at 90 GB/s. Real numbers vary ±30% with runtime, drivers, and batch settings.