The short answer: answering a user means multiplying billions of numbers for every word, and a graphics processing unit (GPU) is built to do exactly that kind of simple, repeated arithmetic in parallel. The longer answer is about memory. To produce one word the hardware has to move the whole model from memory into its processors, and ordinary computer memory is too narrow a channel for that. Data-centre GPUs sit next to very wide memory that can deliver terabytes a second.
How much arithmetic goes into one word
A model does not read text as people do. The application splits your message into tokens, fragments of words, and turns each token into a vector, a long list of numbers that stands for its meaning. The model’s knowledge lives in its weights, arranged in large grids called matrices.
Writing the reply means multiplying that vector through the matrices of one layer after another until a new token comes out. The arithmetic is simple and enormous: about two operations, a multiplication and an addition, for every parameter, for every token. There is no shortcut and no lookup table; every weight has to be used.
| Model size | Operations for one token |
|---|---|
| 8 billion parameters | about 16 billion |
A 500-word reply therefore runs to trillions of operations in a few seconds. On top of that, the model relates each new token to every earlier token in the conversation, work that grows with the square of the conversation’s length.
Few clever cores against thousands of simple ones
A central processing unit (CPU) is built for long, unpredictable chains of logic where step two waits on step one. Much of its silicon goes to predicting branches, reordering work and large caches, so it only has room for a handful of cores. Matrix multiplication does not need any of that: multiplying the top-left number of a grid has nothing to do with the bottom-right one. It needs many basic calculators working at once, which is what a GPU provides.
| Feature | High-end desktop CPU | Data-centre GPU |
|---|---|---|
| Cores | 16 | 16,896 |
| Built for | Complex steps in sequence | Simple arithmetic in parallel |
| Branch prediction | Extensive | Minimal |
| Fit for running AI models | Poor | Excellent |
GPU cores work in step: one instruction is issued and thousands of cores apply it to thousands of numbers at the same moment. When a vector meets a weight matrix, a GPU gets through the layer far sooner than a CPU could.
The memory wall sets the speed limit
Thousands of cores help only if data reaches them fast enough. Their own caches are far too small to hold a model, so the matrices live in the card’s video memory and the entire model has to travel to the cores for every token. The arithmetic is nearly instant; the wait is for the numbers. That delivery rate is memory bandwidth, measured in gigabytes per second, and the fastest a model can possibly produce a token is its size in bytes divided by that bandwidth. No software tuning gets under that floor.
Size follows precision: four bytes per parameter at 32 bits, two at 16 bits, one at 8 bits and about half a byte at 4 bits. An 8-billion-parameter model takes about 16 GB at 16 bits before any conversation is added, and about 4.5 to 5 GB at 4 bits.
| Hardware | Memory bandwidth | Floor per token for a 16 GB model |
|---|---|---|
| Fast desktop, two channels of current desktop memory | about 96 GB/s | about 0.166 seconds |
| Data-centre GPU with stacked high-bandwidth memory | 3,350 GB/s | about 0.004 seconds |
The desktop figure assumes the processor spends no time on the arithmetic, which it cannot, so real speed is lower still. The stacked memory beside a data-centre GPU is what lets it write faster than a person reads.
Why answering one user wastes a GPU, and how providers fix it
Answering a single question is a batch of one. The card loads the 16 GB model, computes one token, and the cores then sit idle until the next transfer arrives. This is called being bandwidth-bound, and it is the normal state when generating text.
Training is the opposite. A layer is loaded once and thousands of examples pass through it, so the cores stay busy long after the transfer; training is compute-bound. Hosted providers make answering look more like training with continuous batching:
- Collect requests from, say, fifty users over a few milliseconds.
- Load the model weights into the cores once.
- Multiply all fifty input vectors in parallel.
- Return one new token to each of the fifty users.
One memory read now serves fifty tokens instead of one, which is how providers get full use from their hardware.
The memory a conversation takes up
The weights are not the only thing in video memory. To avoid recomputing the whole conversation for each new word, the model keeps the working state of every earlier token in a store called the key-value cache. The weights are shared by every user; the cache belongs to one conversation, so fifty batched users mean fifty caches held at once. A long document pasted into a chat makes the cache grow in step with its length, and when video memory runs out the system fails with an out-of-memory error. Engineers trade model size against the longest conversation they allow.
- Fixed memory
- Holds the model’s weights and does not change while it writes.
- Growing memory
- Holds the key-value cache and grows with every token added to the conversation.
A 32 GB gaming card holding a 16 GB model has 16 GB left for caches. Serve several users or one long document and that space is gone. Data-centre cards with 80 GB or more leave room for long business conversations.
Where gaming graphics cards fall short
Gaming cards have fast memory but not much of it: current high-end ones carry 24 GB or 32 GB, checked 5 September 2026. An 8-billion-parameter model at 16 bits (16 GB) fits; larger models do not, unless they are compressed to 4 bits, which costs reasoning quality (see the note on quantization).
- Fast memory for small models
- Enough memory for an 8-billion-parameter model
- Capacity for large models
- Fast links to pool memory across several cards
A model that needs 140 GB has to be spread over several cards. Gaming cards talk to each other over the motherboard’s ordinary expansion slots, too narrow for that job, while data-centre cards carry 80 GB each and dedicated links that move data between cards at hundreds of gigabytes a second. Gaming cards are also designed for a few hours of play, not the constant heat of serving customers day and night; their cooling, drivers and warranties reflect that.
Renting capacity against owning hardware
For most businesses the hardware question ends with not buying any. Data-centre cards are a large capital outlay and lose value fast as new designs arrive, and owning them brings power, cooling, rack space and engineers to keep the drivers running. Owning makes sense only with a steady, heavy volume of work or tight privacy obligations, and few small or mid-size teams are in that position.
A hosted service rents you a slice of the right hardware, charged by the tokens you use. Because providers batch many users together, the running cost that reaches a business is modest. We cost each conversation at about 2,000 tokens in and 500 out, measure it in production rather than from benchmark tables, and route easy turns to a small quick model and hard ones to a larger model. Nobody asking for your opening hours needs the largest one.
What GPUs mean when you choose an AI system
A GPU only calculates the likeliest next word. It holds no business rules, does not know your stock and has no stake in the truth; it will produce a wrong answer as fast as a right one. When an assistant quotes the wrong figure, the hardware did not fail. The system around the model failed to give it the right context or hold the right limits.
1. Adversarial test of our own public assistant, 4 Sep 2026. The full account is in the case study.
Measured once, on our own assistant. A record of that run, not a promise for yours.
Almost none of those faults came from the model. A reliable assistant depends on the instructions, documents and routing that feed the hardware, and you do not need to own the hardware to control those. The agents we build run on hosted capacity; if you would like to know what yours would need, describe it in writing.