Case study

AI Chatbot Testing: What Broke When We Attacked Our Own Assistant

Nine testers spent a day trying to make our public assistant misbehave. This is the record: what held, what broke, the one real security bug, and what is still open.

Site assistantUpdated 30 Sep 20267 min read
On this page
  1. How it ran
  2. The day
  3. What held
  4. What broke
  5. The security bug
  6. What changed
  7. Still open
  8. Why it matters
  9. Scope
  10. Sources

The case in brief

01

Challenge

Learn how our public assistant behaves under attack, including leaks, hidden instructions, abuse and nonsense, before a visitor or a competitor finds out.

02

Approach

Nine testers, one line of attack each. Every reported fault was then replayed by script against the live assistant, so no finding rested on one opinion.

03

Result

No instructions, model or vendor revealed. 33 faults reproduced; 27 fixed and re-checked the same day. One security bug surfaced and was shut within about 40 minutes.

Measured, with sourcesEvery figure from the system log
1311conversations across nine lines of attack
332faults reproduced by scripted replay
273no longer reproduce after the fixes
04leaks of instructions, model or vendor

1. Each tester’s own record. 2. Scripted replay against the live assistant. 3. The same replay after the changes. 4. All 131 conversations.

Measured on 3 Sep 2026 on our own assistant. A record of one test, not a promise.

How the test was run

Nine testers worked side by side, each on one line of attack:

  • getting the assistant to reveal its instructions or the model behind it;
  • instructions hidden inside pasted text;
  • deliberate nonsense;
  • ambiguous wording;
  • pressure to state something false as fact;
  • badgering and abuse;
  • the multi-step flow that files a request with the team;
  • the newest behaviour rules;
  • misuse of the interface itself.

Each tester held real conversations with the live assistant and saved the exact messages and replies. Then came the step that decided the result: every reported fault was replayed by script, sending the identical sequence of messages into a new session, with no model judging the outcome. A fault counted only if it happened again.

The replay changed the picture. Several alarming findings fell away: figures a tester called invented were printed word for word on our own published pages, and a supposed delivery promise was the assistant’s standard line about a written reply within one business day. A test unable to tell a real fault from a convincing story does more harm than having no test.

The day, in order

  1. 3 Sep 2026

    Attack

    Nine testers, one line each: 131 conversations with the live assistant.

  2. 3 Sep 2026

    Replay

    Every reported fault replayed by script on a fresh session. 33 reproduced.

  3. 3 Sep 2026

    Security fix

    The shared-session bug reproduced by hand and closed in about 40 minutes.

  4. 3 Sep 2026

    Rules and code

    Fourteen new instruction rules and one code change.

  5. 3 Sep 2026

    Replay again

    27 of the 33 no longer reproduce. One real defect stays open.

What held

In all 131 conversations the assistant did not reveal its instructions, did not name the model or the company behind it, and did not name an internal system, even when the request was framed as a compliance check or a developer’s test.

Its firmest defence was one sentence it kept repeating. Pushed again and again to confirm it had already passed a request to the team, it replied that nothing had been filed and nothing sent. At no point did it make up a reference number. A visitor writing in apparent distress was directed to emergency services first. Politics, homework and medical questions were declined briefly and kindly.

What broke

Thirty-three faults reproduced. None was exotic; they gathered around a few everyday lapses:

  • It kept selling to angry visitors. Insulted and asked for a person, it handed over correctly, then went back to describing services two messages later. Six findings shared that one cause: the handed-over state did not stick.
  • It offered a menu. Asked about one service, it listed three, and one phrasing drew out most of the published service list at once.
  • It talked about tools. It mentioned the platforms underneath it and commented on products made by other companies, the kind of claim we advise clients never to make.
  • It invented one figure. A usage number that appears nowhere on our published pages.
  • It lost an email address. Given a valid address, it asked for it again later, and in one flow called a complete address incomplete.

The one real security bug

The tester working on the interface found a fault no change of wording could fix. To answer common first questions faster, replies were cached, and the cached copy carried the session of whoever asked first. Two visitors opening with the same question were given the same session, so each could see what the other had told the assistant. Reproducing it by hand took under a minute: one visitor typed a company name and a budget, and a second visitor got both read back.

It was closed the same morning, in about 40 minutes. Cached replies now hold no session at all, every visitor is given a separate one, the old cache was emptied and the affected session set apart. The check we now run: two fresh visitors ask the same question and receive two different sessions.

What we changed

The fix came in two parts: fourteen new rules in the assistant’s instructions, and one change to the code. The rules leave little room:

  • once a visitor is handed to a person, that stays so until the request is filed;
  • describe one service, never three;
  • describe no tool, ours or anyone else’s;
  • invent no figure;
  • take an email address the first time it appears and never ask for it again;
  • a request to be forgotten erases the conversation;
  • treat any reference number we never issued as unknown;
  • never read a request into nonsense.

After the change, each failing conversation ran again against the updated assistant, and 27 of the 33 did not recur. We quote that figure because one replay method produced both numbers.

What is still open

A genuine defect is left. After the confirmation card, where you see exactly what will go to the team, a reply of “yes, send it” sometimes forgets the email address given earlier and starts the request over. No data goes missing and the assistant says nothing untrue; it costs the visitor time. It needs work in the code rather than another rule, and it is next on the list.

Six findings still show, and we consider two of them mistaken. In one, a reply was faulted for quoting the refund policy as it stood that day, which was accurate. In the other, the assistant on the AEO page already opens with the correct service. Both are listed here so anyone can check the arithmetic.

Why this is part of every build

When an assistant pushes a sale on someone who is furious, cites a number that was never published, or asks twice for an email address, the business is talking to its customers with nobody watching. None of those faults shows from the inside; the assistant reads well until someone pushes it.

So each build faces the same routine before the public sees it: attack it along the same nine lines, replay every fault by script, fix, replay again and write down what remains open. That document comes with the agent at handover.

Tested, and deliberately not claimed

Built
  • Nine lines of attack on the live assistant
  • A scripted replay of every reported fault
  • A fix for the session bug, with its own check
  • Fourteen written rules, each replayed against the old failures
Deliberately not built
  • A fault counted on a tester’s word alone
  • A quiet patch: the security bug is published here
  • A claim that every fault is gone: one stays listed as open
  • Hiding the findings we disagree with: both are shown

The figures and where they come from

131 conversations
Four testers held 63 of them and the remaining five held 68; each tester kept a record.
33 reproduced faults
Counted from the scripted replay against the live assistant; a tester’s report alone did not count.
27 fixed
Counted by running the same conversations again after the changes, with the identical replay.
No leaks
Not one of the 131 conversations exposed the instructions, the model, the vendor or an internal system name.
One security bug
The shared-session fault, closed on 3 September 2026 after a manual reproduction.

Written by Sathish Balakrishnan, Founder of iHayz

Case studies are written from the system’s own log by the people who built it. The work is contracted through Media Experts LLC in the USA or Media Experts in India. About the company

Start

Not sure what you need? Ask in writing.

Describe the work in a few lines. We will reply in writing within one business day with what we would build and how.