Notes

AI Model Quantization Explained: Smaller Models and What They Lose

Quantization makes a model smaller by rounding its learned numbers. The same 8-billion-parameter model needs 32 GB at full precision and under 5 GB at 4 bits; the saving is large, and the loss appears in predictable places first.

GuideUpdated 30 Sep 20266 min read
On this page
  1. Precision of one weight
  2. Smaller is also faster
  3. Memory arithmetic
  4. Outlier weights
  5. Where rounding hurts
  6. How models get compressed
  7. What it means for you

A model is, underneath, a very long list of weights stored at high precision. Quantization stores them in smaller formats, so the model runs in far less memory, which lets large systems fit on modest hardware. The price is precision: everyday chat hardly notices, while demanding reasoning becomes measurably weaker.

What precision means for a single weight

Weights are normally kept as floating-point numbers, which carry several significant digits and a movable decimal point so that very small fractions can be stored exactly. Quantization maps them onto a fixed grid of integer levels instead.

FormatTypeFraction bitsDistinct levels
32-bitFloating point23about 4.3 billion
16-bitFloating point1065,536
8-bitIntegernone256
4-bitIntegernone16

Follow one weight through the formats and you can watch the detail go:

FormatPrecision keptThe same weight
32-bitabout 7 significant digits0.0731000
16-bitabout 3 to 4 significant digits0.0731
8-bit256 levels0.075
4-bit16 levels0.07

Training adjusts each connection by tiny amounts, which is why it needs the fine resolution. Once weights are snapped to the nearest notch on the grid, each connection is a little off from what training decided, and the model loses a little of its finesse.

Smaller weights also mean quicker replies

The processor is seldom what slows a model down. The limit is how quickly weights travel from the memory chips to the processor, because producing each word means reading the whole list of weights once more.

Think of a checkout with a fast cashier and a narrow conveyor belt: the cashier spends most of the time waiting for goods to arrive. Language models do very little arithmetic per weight but must touch every weight for each token, so they are always waiting on the belt. Shrinking the weights widens it. Unpacking the compressed numbers costs far less time than the transfer it saves, which is the main reason models can run on local machines at all; without compression, ordinary hardware would produce text at a crawl.

How much memory each format needs

The memory needed is a floor, not an estimate: bytes per parameter times the number of parameters. If the compressed weights are larger than the memory available, the model does not load. The same model can overwhelm a server at full precision and run on a laptop once heavily compressed.

Arithmetic checked 5 Sep 2026
FormatBytes per parameterMemory for an 8-billion-parameter model
32-bit432 GB
16-bit216 GB
8-bit18 GB
4-bitabout 0.54.5 to 5 GB

A compressed file is slightly larger than the bare arithmetic suggests, because it carries scaling factors that let the rounded weights be read back correctly. Even so, going from 32 GB to about 4.5 GB is a very large saving.

Why a few large weights carry so much

In a big network most weights sit close to zero and stand for weak links between ideas; rounding them does little harm. A small share grow unusually large during training, and those tend to carry rigid behaviour such as arithmetic, formatting and strict logical steps.

A single grid stretched over the whole model flattens these outliers to its top notch, and they lose their scale relative to everything else. The result is a model that writes fluent, confident text with broken logic inside it. Current methods split the weights into small blocks and give each block its own scale, so a block with an outlier can stretch its grid without blurring the ordinary weights around it.

Which tasks lose quality first

The damage is uneven. For most models, moderate rounding costs almost nothing, and aggressive rounding still copes with general chat, summaries and pulling facts out of text. Past that point quality falls away quickly.

TaskEffect of heavy compressionWhy
General chatSmallLanguage repeats itself, so small errors are absorbed
SummariesSmallThe main meaning survives rounding
Multi-step reasoningSevereErrors add up along the chain
Exact arithmeticSevereIt leans on precise outlier weights
Writing codeSevereOne wrong character stops the program

Rare words go early too. A single mistaken variable name breaks a script, and one faulty step spoils a reasoned conclusion, while a slightly off word in a chat reply usually goes unnoticed.

How the rounding is applied, and the files you meet

MethodWhenComputing neededQuality kept
Post-training quantizationAfter training endsLowGood
Quantization-aware trainingDuring trainingHighExcellent

Post-training quantization takes a finished model, uses a small calibration set to find the weights that matter most, and rounds the rest accordingly, in minutes or hours. Quantization-aware training builds the rounding into training so the model learns to live with it, which keeps more quality but needs the resources of a full training run. Nearly every compressed file people download was made the first way: model makers usually publish the full-precision release, and the open community calibrates and publishes compressed versions, often dozens of them within hours.

The files fall into two families. One is a single-file package for ordinary processors and mixed processor-and-graphics-card setups, holding the weights, the tokeniser and the settings together; its file names encode the base precision, whether sensitive layers were kept at higher precision, and where the size-to-quality balance sits. The other family stores weights only, laid out for the wide, parallel way a graphics card reads memory, and keeps a small share of the most important weights at higher precision so the card can unpack them without stalling. We checked both families against their project documentation on 5 September 2026.

What quantization means when you choose an AI system

For a business the one question is whether answers to your customers stay correct after compression, and the only way to know is to test it on your own real questions. Hosted services compress and tune their models behind the scenes too, without publishing the details, so the same advice applies there.

Most of what customers ask a business, such as a refund policy, opening hours or a delivery area, needs a retrieved fact rather than long reasoning, and that is where a heavily compressed small model does well. We send simple turns to a small quick model and harder ones to a larger model, and that routing runs in production. When we tested our own public assistant, the faults a customer would notice came from rules, document retrieval and handover; the precision of the weights played no part.

1311conversations tested
331faults reproduced
271faults fixed
141rules added
11code defect corrected

1. Adversarial test of our own public assistant, 4 Sep 2026. The full account is in the case study.

Measured once, on our own assistant. A record of that run, not a promise for yours.

If choosing file formats and hardware is not how you want to spend your time, a hosted service takes the whole question away. Our knowledge assistants run on hosted models, compressed where that is safe and at full strength where the question needs it.

Start

Not sure what you need? Ask in writing.

Describe the work in a few lines. We will reply in writing within one business day with what we would build and how.