Notes

What Are AI Parameters, and Does the Count Matter?

A parameter is a number a model learned in training, either a weight or a bias. The count sets how much memory the model needs; it says very little about whether the model will do your job well.

GuideUpdated 30 Sep 20268 min read
On this page
  1. What a parameter is
  2. Layers and connections
  3. How the numbers are set
  4. Memory arithmetic
  5. Size is not quality
  6. Total vs active
  7. Open weights and licences
  8. What it means for you

Model announcements lead with a parameter count, as if it were a score. It is closer to a size on a shipping label: it tells you how much space and power the thing needs, and nothing about how well it will answer your customers.

A parameter is a learned number

Underneath the talk of digital brains, a model is a long chain of ordinary arithmetic. It is made of small junctions called nodes. A number arrives at a node, is multiplied, has something added, and moves on to the next layer. The numbers doing the multiplying and the adding are the parameters, and there are two kinds.

A weight works like a volume control on one connection: it sets how much an incoming signal counts. In a system predicting whether a customer will ask for a refund, “the item arrived broken” would carry a heavy weight and “it is Tuesday” almost none. A bias is a starting level added after the weights, the default before any input is counted. Even a shop with perfect shipping sees a few returns a day, and the bias keeps the prediction from ever falling to zero.

So a seven-billion-parameter model contains seven billion of these weights and biases, stored as a fixed table. Your words become numbers, pass through that table, and the numbers that come out are turned back into text.

Parameter
A number the model learned; each one is a weight or a bias.
Weight
Sets, by multiplying, how much one incoming signal matters.
Bias
A starting value added to a calculation to move the result up or down.
Node
One junction in the network where weights and a bias are applied.

Why the count climbs into billions

Nodes sit in layers: an input layer that takes the prompt, an output layer that produces the reply, and dozens of hidden layers between them. Each node connects to every node in the next layer, and each connection has its own weight. Two neighbouring layers of a thousand nodes each need a million weights between them. Current language models stack dozens of layers, each with thousands of nodes, so the total reaches billions quickly.

The layers handle language at different depths. Early ones pick up grammar, such as telling a noun from a verb. Middle ones join words into ideas, for instance that apples and oranges are both fruit. Late ones shape those ideas into a reply. None of it is thought in the human sense; it is statistical filtering, and more parameters let the filter hold finer patterns of language.

Training sets the numbers; answering only reads them

A new model starts with every weight and bias set at random, and at that stage it answers in gibberish. Training turns those random numbers into something useful by a long loop:

  1. Predict the next token of a piece of text using the current numbers.
  2. Measure the gap between the guess and the real next token.
  3. Work out how much each parameter contributed to the error.
  4. Nudge every weight and bias a tiny amount in the helpful direction.
  5. Repeat, billions of times, over trillions of words.

The method is called gradient descent. Picture walking down a hill in thick fog: you cannot see the valley, but you can feel which way the ground falls and take a step that way, again and again, until the error is low.

When training ends, the numbers are frozen. Using the model to answer, called inference, only reads them. That is why an assistant does not learn a new fact because you told it one in a chat; what it knows about your business has to be handed to it with each question.

The count decides the memory bill

Every parameter has to sit in the memory of the graphics card, its video memory, while the model runs. If the model does not fit, it either refuses to load or runs too slowly to use. How much room each parameter takes depends on its precision: 32 bits is four bytes, 16 bits is two, and the compression called quantization brings it to one byte at 8 bits or half a byte at 4 bits, giving up a little nuance for a large saving.

Weights only, before the conversation’s own memory and other overhead
PrecisionBytes per parameterMemory for a 7-billion-parameter model
32-bit4about 28 GB
16-bit2about 14 GB
8-bit1about 7 GB
4-bit0.5about 3.5 GB

A 70-billion-parameter model at 16-bit needs 140 GB just to be loaded, which means several data-centre graphics cards linked together. That is why most businesses reach models through a hosted service. The note on quantization covers what the rounding costs.

Base models, instruction tuning and why smaller can win

More parameters give a model more room to store facts and patterns. They do not tell you whether it will handle your task well.

A model straight out of its first training is a base model, and it only continues text. Give it “What is your return policy?” and it may reply “What is your shipping policy?”, taking the prompt for the start of a list. It has billions of parameters and no idea how to be an assistant.

A second stage, instruction tuning, teaches it to answer questions, keep to a format and stay polite. The quality of that stage counts for more than raw size, and a small model with good tuning beats a large base model. In production we regularly see a newer, smaller model follow strict rules and extract data better than an older, larger one. Model makers now put their effort into better data and tuning for smaller networks rather than into sheer size.

Total parameters and active parameters

The headline figure can mislead. Older models were dense: every parameter took part in producing every token. Many large models today use a mixture-of-experts design, several smaller networks under one roof with a router that sends each input only to the experts suited to it. The rest stay idle for that step.

DesignParameters heldParameters used per tokenCompute per token
DenseAllAllHigh
Mixture of expertsAllA fractionMedium

The total sets how much memory you need to load the model. The active share sets how much computing you pay for on each token. A very large advertised count usually refers to the total of a mixture-of-experts model, which matters when you estimate hosting or hardware.

Downloadable weights come with a licence

Some businesses want to run a model on their own machines to keep customer data in-house. Downloading an open model means downloading the exact weights and biases its maker trained. Having the numbers is not the same as being free to use them as you like; the licence decides that.

When we checked the licences of several open models against their source files on 4 September 2026, the terms differed widely. One asked for its name to be shown in the product, but only above a very large user or revenue threshold. Another carried a standard permissive licence with no threshold and no attribution. A third allowed commercial use but required a separate agreement above a very large number of monthly users.

What the count means when you choose an AI system

Our own assistants are judged in production on three things that touch the business: what one conversation costs to run, how it refuses, and how long the reply takes. For costing we treat a conversation as about 2,000 tokens in and 500 tokens out. The parameter count of the largest model has no bearing on a question like “what time do you open”, so simple turns go to a small, quick model and harder ones, such as a warranty claim on a broken delivery, to a larger one.

When something goes wrong, retraining (changing the parameters) is slow, costly and unpredictable, and it is seldom needed. When we tested our own public assistant, almost every fault came from the instructions or the documents, not the model.

1311test conversations
331faults reproduced
271fixed with 14 new rules and one code fix
01leaks of instructions or vendor

1. Adversarial test of our own public assistant, 4 Sep 2026; the model’s parameters were unchanged throughout. Details in the case study.

Measured once, on our own assistant. A record of that run, not a promise for yours.

Built
  • Clear formatting and refusal rules in the instructions
  • Accurate documents for the model to read
  • Complex or sensitive questions routed to a person
Deliberately not built
  • Relying on what the model memorised for business facts
  • Fine-tuning the parameters to fix a small habit
  • Assuming a larger model stops inventing things

Your customers will never ask how many parameters the assistant has. They want a fast, correct answer and a person when it matters. Choose the system on those terms, and on a running cost that fits your margins; knowledge assistants are built that way.

Start

Not sure what you need? Ask in writing.

Describe the work in a few lines. We will reply in writing within one business day with what we would build and how.