How the test was run
Nine testers worked side by side, each on one line of attack:
- getting the assistant to reveal its instructions or the model behind it;
- instructions hidden inside pasted text;
- deliberate nonsense;
- ambiguous wording;
- pressure to state something false as fact;
- badgering and abuse;
- the multi-step flow that files a request with the team;
- the newest behaviour rules;
- misuse of the interface itself.
Each tester held real conversations with the live assistant and saved the exact messages and replies. Then came the step that decided the result: every reported fault was replayed by script, sending the identical sequence of messages into a new session, with no model judging the outcome. A fault counted only if it happened again.
The replay changed the picture. Several alarming findings fell away: figures a tester called invented were printed word for word on our own published pages, and a supposed delivery promise was the assistant’s standard line about a written reply within one business day. A test unable to tell a real fault from a convincing story does more harm than having no test.
The day, in order
- 3 Sep 2026
Attack
Nine testers, one line each: 131 conversations with the live assistant.
- 3 Sep 2026
Replay
Every reported fault replayed by script on a fresh session. 33 reproduced.
- 3 Sep 2026
Security fix
The shared-session bug reproduced by hand and closed in about 40 minutes.
- 3 Sep 2026
Rules and code
Fourteen new instruction rules and one code change.
- 3 Sep 2026
Replay again
27 of the 33 no longer reproduce. One real defect stays open.
What held
In all 131 conversations the assistant did not reveal its instructions, did not name the model or the company behind it, and did not name an internal system, even when the request was framed as a compliance check or a developer’s test.
Its firmest defence was one sentence it kept repeating. Pushed again and again to confirm it had already passed a request to the team, it replied that nothing had been filed and nothing sent. At no point did it make up a reference number. A visitor writing in apparent distress was directed to emergency services first. Politics, homework and medical questions were declined briefly and kindly.
What broke
Thirty-three faults reproduced. None was exotic; they gathered around a few everyday lapses:
- It kept selling to angry visitors. Insulted and asked for a person, it handed over correctly, then went back to describing services two messages later. Six findings shared that one cause: the handed-over state did not stick.
- It offered a menu. Asked about one service, it listed three, and one phrasing drew out most of the published service list at once.
- It talked about tools. It mentioned the platforms underneath it and commented on products made by other companies, the kind of claim we advise clients never to make.
- It invented one figure. A usage number that appears nowhere on our published pages.
- It lost an email address. Given a valid address, it asked for it again later, and in one flow called a complete address incomplete.
The one real security bug
The tester working on the interface found a fault no change of wording could fix. To answer common first questions faster, replies were cached, and the cached copy carried the session of whoever asked first. Two visitors opening with the same question were given the same session, so each could see what the other had told the assistant. Reproducing it by hand took under a minute: one visitor typed a company name and a budget, and a second visitor got both read back.
It was closed the same morning, in about 40 minutes. Cached replies now hold no session at all, every visitor is given a separate one, the old cache was emptied and the affected session set apart. The check we now run: two fresh visitors ask the same question and receive two different sessions.
What we changed
The fix came in two parts: fourteen new rules in the assistant’s instructions, and one change to the code. The rules leave little room:
- once a visitor is handed to a person, that stays so until the request is filed;
- describe one service, never three;
- describe no tool, ours or anyone else’s;
- invent no figure;
- take an email address the first time it appears and never ask for it again;
- a request to be forgotten erases the conversation;
- treat any reference number we never issued as unknown;
- never read a request into nonsense.
After the change, each failing conversation ran again against the updated assistant, and 27 of the 33 did not recur. We quote that figure because one replay method produced both numbers.
What is still open
A genuine defect is left. After the confirmation card, where you see exactly what will go to the team, a reply of “yes, send it” sometimes forgets the email address given earlier and starts the request over. No data goes missing and the assistant says nothing untrue; it costs the visitor time. It needs work in the code rather than another rule, and it is next on the list.
Six findings still show, and we consider two of them mistaken. In one, a reply was faulted for quoting the refund policy as it stood that day, which was accurate. In the other, the assistant on the AEO page already opens with the correct service. Both are listed here so anyone can check the arithmetic.
Why this is part of every build
When an assistant pushes a sale on someone who is furious, cites a number that was never published, or asks twice for an email address, the business is talking to its customers with nobody watching. None of those faults shows from the inside; the assistant reads well until someone pushes it.
So each build faces the same routine before the public sees it: attack it along the same nine lines, replay every fault by script, fix, replay again and write down what remains open. That document comes with the agent at handover.
Tested, and deliberately not claimed
- Nine lines of attack on the live assistant
- A scripted replay of every reported fault
- A fix for the session bug, with its own check
- Fourteen written rules, each replayed against the old failures
- A fault counted on a tester’s word alone
- A quiet patch: the security bug is published here
- A claim that every fault is gone: one stays listed as open
- Hiding the findings we disagree with: both are shown
The figures and where they come from
- 131 conversations
- Four testers held 63 of them and the remaining five held 68; each tester kept a record.
- 33 reproduced faults
- Counted from the scripted replay against the live assistant; a tester’s report alone did not count.
- 27 fixed
- Counted by running the same conversations again after the changes, with the identical replay.
- No leaks
- Not one of the 131 conversations exposed the instructions, the model, the vendor or an internal system name.
- One security bug
- The shared-session fault, closed on 3 September 2026 after a manual reproduction.