Rankings of AI models reshuffle every few months, so a guide that crowns one product is out of date before the year ends. What stays useful is the set of differences a customer feels: how fast the reply comes, what each conversation costs you to run, and what the system does once the answer is not in front of it.
Those differences depend far more on how the chatbot is built than on whose model sits underneath. We looked at the field against published documentation on 4 September 2026, and we build and run customer-facing assistants on more than one family of models, so the notes below come from running them rather than from a leaderboard.
Four kinds of chatbot side by side
| Type | What it does | Who sets it up | When it does not know | Who holds the data |
|---|---|---|---|---|
| Rule-based menu bot | Shows buttons and fixed replies written in advance | You, in the tool’s own editor | Sends the visitor back to the menu or to a contact form | The company that hosts the flows |
| Hosted subscription chatbot | A general-purpose model reads your site or a pasted document and writes free-text replies | You, through the tool’s settings | Unless told otherwise, it tends to produce a believable guess | The tool’s provider, under its terms |
| Assistant that answers from your documents | Looks up the documents you approved and replies from them, naming the source | A builder, from your policies and answers | Says it does not know and passes the visitor to a person | You, when it runs on accounts in your name |
| Agent that acts in your systems | Works inside your booking, inbox or order tools: books, drafts, updates records | A builder, with approval gates you decide | Stops and asks a person before any step it is not cleared for | You, inside your own systems |
No row is right for everyone. A shop that gets the same five questions all week needs something different from a clinic whose messages turn into bookings. The sections below take each type in turn.
Hosted subscription chatbots: quick to start, loose at the edges
A subscription chatbot tool connects a large hosted model to your website text and is live within a day. For general questions the replies read well. The weak spot is the edge of what it was given: asked about a delivery charge or a refund rule it has no document for, a general-purpose model will often write a confident answer that nobody at the business ever approved.
Two things to check before signing up. First, whether the tool can be told to refuse and hand over rather than guess. Second, what happens to your customers’ conversations: builders who work for businesses use service agreements under which the model provider may not use your conversation data for training. Ask the tool for the same assurance in writing.
Assistants that answer from your own documents
This type is often described as a chatbot “trained on your data”, though the model itself is not retrained. Before each reply, the assistant fetches the passages from your approved documents that match the question and answers from those, so a change to a policy takes effect as soon as the document changes. Our knowledge assistants work this way.
How much it can read at once is set by its context window, the working memory for one conversation. Current windows reach around a million tokens, roughly a dozen thick books, far more than any support chat uses. Room to read everything is not a reason to send everything: each piece of text the model reads is billed, so the documents should be the right ones, not all of them.
Agents that take actions in your systems
An agent goes a step further than answering. It can book an appointment, draft a reply in your inbox or update an order record. That makes it the most useful type and the one that needs the firmest limits. The AI agents we build, including the Inbox Agent, work within approval gates you set in writing.
Why one model for every message is a poor fit
Model families come in sizes. Small models reply fast and cost little to run per conversation; large ones reason better but take longer to write each reply and cost more. Telling a customer when you open does not call for the largest model, and a visitor who waits several seconds for that answer tends to assume the widget has broken.
A well-built assistant therefore routes each message:
- A small classifier reads the incoming message. It writes nothing; it only sorts.
- Routine turns such as greetings, order numbers and opening hours go to a small, quick model.
- Harder turns such as a refund rule or a disputed charge go to a larger, more careful model.
The switch takes milliseconds and the customer never sees it. Most replies arrive at once, and the higher running cost applies only to the few messages that need it. Running cost is also the smaller part of the bill: building the system, keeping its rules current and hosting it is where the work sits, and all of it belongs in a fixed price in the written proposal.
What went wrong when we tested our own assistant
We attacked our own public assistant in early September 2026, with testers playing angry customers, chasing refunds and trying to pull out its hidden instructions.
1. Adversarial test of our own public assistant, 4 Sep 2026. The full account is in the case study.
Measured once, on our own assistant. A record of that run, not a promise for yours.
Several alarming findings fell away on replay; prices an attacker called invented turned out to be printed on our own published sheet. The faults that held up were mostly missing rules, not a weak model:
- Pushing a sale on a furious customer when its instructions said to bring in a person.
- Dumping the whole menu instead of the one service asked about.
- Oversharing by naming the tools the build used.
- One invented usage figure that no document contained.
- Asking a second time for an email address that was already in the chat.
Twenty-seven of the faults were fixed by adding explicit rules, with no change of model. Across every conversation nothing leaked: no system instructions, no model name, no vendor. The one serious bug was outside the AI altogether. A caching layer stored the first greeting together with the session of whoever triggered it, so two visitors could share one conversation memory. It was reproduced by hand and closed the same morning. The lesson for a buyer is that the code around the model deserves as much scrutiny as the model.
How to test a chatbot before customers see it
Chatting with a widget for ten minutes proves little, and asking a model to grade its own replies is a guess. A test you can repeat takes an afternoon:
- Step 1
Define the attacks
Write down the ways a customer could push the assistant past its rules: anger, refunds, requests for its instructions.
- Step 2
Run them live
Use the real widget, not a simulator, so the whole system is tested.
- Step 3
Record everything
Keep every input and every output exactly as it happened.
- Step 4
Replay mechanically
Feed the same inputs back with a script. A fault that returns is a bug to fix; one that never returns was noise. Fix, then replay again.
Models can answer the same question differently from one run to the next, which is why replay matters. While testing, watch for a polite refusal on out-of-scope questions and a clean handover when the customer grows frustrated.
A builder worth hiring tells you what the system cannot do. If you want that answer for your own case, describe the job in writing.