A model is, underneath, a very long list of weights stored at high precision. Quantization stores them in smaller formats, so the model runs in far less memory, which lets large systems fit on modest hardware. The price is precision: everyday chat hardly notices, while demanding reasoning becomes measurably weaker.
What precision means for a single weight
Weights are normally kept as floating-point numbers, which carry several significant digits and a movable decimal point so that very small fractions can be stored exactly. Quantization maps them onto a fixed grid of integer levels instead.
| Format | Type | Fraction bits | Distinct levels |
|---|---|---|---|
| 32-bit | Floating point | 23 | about 4.3 billion |
| 16-bit | Floating point | 10 | 65,536 |
| 8-bit | Integer | none | 256 |
| 4-bit | Integer | none | 16 |
Follow one weight through the formats and you can watch the detail go:
| Format | Precision kept | The same weight |
|---|---|---|
| 32-bit | about 7 significant digits | 0.0731000 |
| 16-bit | about 3 to 4 significant digits | 0.0731 |
| 8-bit | 256 levels | 0.075 |
| 4-bit | 16 levels | 0.07 |
Training adjusts each connection by tiny amounts, which is why it needs the fine resolution. Once weights are snapped to the nearest notch on the grid, each connection is a little off from what training decided, and the model loses a little of its finesse.
Smaller weights also mean quicker replies
The processor is seldom what slows a model down. The limit is how quickly weights travel from the memory chips to the processor, because producing each word means reading the whole list of weights once more.
Think of a checkout with a fast cashier and a narrow conveyor belt: the cashier spends most of the time waiting for goods to arrive. Language models do very little arithmetic per weight but must touch every weight for each token, so they are always waiting on the belt. Shrinking the weights widens it. Unpacking the compressed numbers costs far less time than the transfer it saves, which is the main reason models can run on local machines at all; without compression, ordinary hardware would produce text at a crawl.
How much memory each format needs
The memory needed is a floor, not an estimate: bytes per parameter times the number of parameters. If the compressed weights are larger than the memory available, the model does not load. The same model can overwhelm a server at full precision and run on a laptop once heavily compressed.
| Format | Bytes per parameter | Memory for an 8-billion-parameter model |
|---|---|---|
| 32-bit | 4 | 32 GB |
| 16-bit | 2 | 16 GB |
| 8-bit | 1 | 8 GB |
| 4-bit | about 0.5 | 4.5 to 5 GB |
A compressed file is slightly larger than the bare arithmetic suggests, because it carries scaling factors that let the rounded weights be read back correctly. Even so, going from 32 GB to about 4.5 GB is a very large saving.
Why a few large weights carry so much
In a big network most weights sit close to zero and stand for weak links between ideas; rounding them does little harm. A small share grow unusually large during training, and those tend to carry rigid behaviour such as arithmetic, formatting and strict logical steps.
A single grid stretched over the whole model flattens these outliers to its top notch, and they lose their scale relative to everything else. The result is a model that writes fluent, confident text with broken logic inside it. Current methods split the weights into small blocks and give each block its own scale, so a block with an outlier can stretch its grid without blurring the ordinary weights around it.
Which tasks lose quality first
The damage is uneven. For most models, moderate rounding costs almost nothing, and aggressive rounding still copes with general chat, summaries and pulling facts out of text. Past that point quality falls away quickly.
| Task | Effect of heavy compression | Why |
|---|---|---|
| General chat | Small | Language repeats itself, so small errors are absorbed |
| Summaries | Small | The main meaning survives rounding |
| Multi-step reasoning | Severe | Errors add up along the chain |
| Exact arithmetic | Severe | It leans on precise outlier weights |
| Writing code | Severe | One wrong character stops the program |
Rare words go early too. A single mistaken variable name breaks a script, and one faulty step spoils a reasoned conclusion, while a slightly off word in a chat reply usually goes unnoticed.
How the rounding is applied, and the files you meet
| Method | When | Computing needed | Quality kept |
|---|---|---|---|
| Post-training quantization | After training ends | Low | Good |
| Quantization-aware training | During training | High | Excellent |
Post-training quantization takes a finished model, uses a small calibration set to find the weights that matter most, and rounds the rest accordingly, in minutes or hours. Quantization-aware training builds the rounding into training so the model learns to live with it, which keeps more quality but needs the resources of a full training run. Nearly every compressed file people download was made the first way: model makers usually publish the full-precision release, and the open community calibrates and publishes compressed versions, often dozens of them within hours.
The files fall into two families. One is a single-file package for ordinary processors and mixed processor-and-graphics-card setups, holding the weights, the tokeniser and the settings together; its file names encode the base precision, whether sensitive layers were kept at higher precision, and where the size-to-quality balance sits. The other family stores weights only, laid out for the wide, parallel way a graphics card reads memory, and keeps a small share of the most important weights at higher precision so the card can unpack them without stalling. We checked both families against their project documentation on 5 September 2026.
What quantization means when you choose an AI system
For a business the one question is whether answers to your customers stay correct after compression, and the only way to know is to test it on your own real questions. Hosted services compress and tune their models behind the scenes too, without publishing the details, so the same advice applies there.
Most of what customers ask a business, such as a refund policy, opening hours or a delivery area, needs a retrieved fact rather than long reasoning, and that is where a heavily compressed small model does well. We send simple turns to a small quick model and harder ones to a larger model, and that routing runs in production. When we tested our own public assistant, the faults a customer would notice came from rules, document retrieval and handover; the precision of the weights played no part.
1. Adversarial test of our own public assistant, 4 Sep 2026. The full account is in the case study.
Measured once, on our own assistant. A record of that run, not a promise for yours.
If choosing file formats and hardware is not how you want to spend your time, a hosted service takes the whole question away. Our knowledge assistants run on hosted models, compressed where that is safe and at full strength where the question needs it.