← All demos

Quantization = sending the photo over Messenger

The platform compresses your picture so it's faster to send. What arrives still looks like your photo — just smaller. Quantization does the same thing to a model's weights: store each number in fewer bits.

The photo you sent — full precision (F16)
What arrives —

The same thing, on actual weights

What it does to the file

The image is an analogy — models quantize numbers, not pixels — but the trade is identical: fewer bits per value → smaller file, faster loading and generation, and a little damage that grows fast below 4 bits.