Case study

Instagram AI Agent Case Study: Our Own Inbox, Live in Four Days

Before an inbox agent answered anyone else’s customers, it ran on ours. Four days from the first test message to a real reply, with the bugs, the refusal and the log that recorded each step.

Inbox AgentUpdated 30 Sep 20269 min read
On this page
  1. The first rule
  2. How it is built
  3. The answer list
  4. Four days
  5. The first reply
  6. The refusal
  7. The bug
  8. Two flaws
  9. Left for a person
  10. Scope
  11. For a client
  12. Definitions

The case in brief

01

Challenge

Show the whole loop working on our own account first: real messages, real replies and real handovers, with no planted questions.

02

Approach

An answer list in the owner’s voice, handover rules enforced in code, and an allow-list so no follower could receive a test reply.

03

Result

Routine questions got an answer within roughly a minute. A refund request acknowledged, stopped and passed to a person with its context. Two bugs found on our own account.

Measured, with sourcesEvery figure from the system log
61 s1first live reply, arrival to sent
~2 s2median engine decision
3 of 93threads drafted; six were photo-only
Every one4money question handed to a person

1–2. Worker log, day 3, 18:44 to 18:45 UTC. 3. Inbox read with skip counters, day 2. 4. The handover list is enforced in the send step; the model cannot override it.

Measured during this test on our own account. A record of one run, not a promise.

The rule set before any code was written

The account used for the test is genuine. It belongs to a small interior design business, has roughly 3,200 followers and receives real messages from real people. So one decision came before everything else. The agent must have no way to reach a real follower by accident.

It started paused. When it was switched on, it ran behind an allow-list and replied to one account only: our own second profile, which played the customer. Messages from everyone else went through every step, read, decided and drafted, and the draft was held. The owner saw every would-be reply; no follower saw any.

Every new client starts the same way. The agent drafts without sending until the owner has checked the drafts, and going live then means deleting one configuration line.

Four stages between the customer and the owner

There is no dashboard and no new app. Customers keep writing in Instagram, and the owner keeps working from Instagram and email. Between them sit four stages:

  1. Delivery. The platform sends a signed notice of each new message.
  2. Queue. The notice is verified, checked for duplicates and saved before any decision is made.
  3. Decision. The engine answers from approved content, or stops.
  4. Record. Each decision becomes a log row and, when needed, an email to the owner.

The link runs through the platform’s official interface. The owner grants it on the platform’s consent page and can withdraw it whenever they choose. The receiving end does no thinking: it checks the signature, saves the event and acknowledges within milliseconds. The decision runs afterwards, once the message is safely stored, so nothing is dropped while a model is busy.

The answer list is all the agent knows

General knowledge is off limits to the engine. What it may use is the answer list: a short document in the owner’s voice with the questions the business receives and the way it answers them. Ours had seven entries, covering how the no-charge first room plan works, what to send for it, and what a whole-house project involves.

Three limits sit around the list, enforced in code rather than asked for in a prompt. These always reach a person:

  • refunds, payments and invoices: anything touching money, whether paid already or still due;
  • complaints, and any message from a customer who is unhappy;
  • any question whose answer is not written in the approved list.

The third limit is the one people do not expect. An unknown question is a handover, not a chance for the model to help. Over the four days the agent did not invent one fact, because the route from uncertainty to the customer passes through a person by design.

Four days, in order

  1. Day 1

    Delivery proven

    A test message sent at 21:16:12 was verified and queued at 21:16:13. An unsigned request was rejected and a replayed event ignored.

  2. Day 1

    Account linked

    The owner approved the link with a single tap on the platform’s consent page. Earlier messages were replayed for context, and 117 posts synced, the same number the profile shows.

  3. Day 2

    Answers on paper

    Seven test messages, all held while the agent was paused: the refund and the unknown question handed over, the safety question refused, the routine ones answered.

  4. Day 3

    Live, behind the allow-list

    The first real reply sent into a thread, and the first real handover emailed to the owner.

  5. Day 4

    Fixes from live traffic

    Two design flaws found by watching real messages, and corrected the same day.

The first real reply took 61 seconds

On day three, at 18:44 UTC, our stand-in customer asked whether free room plans were on offer. By 18:45:09 the reply was in the thread, taken from the approved list and written in the account’s own lowercase style: the first plan costs nothing; send a photo of the room, rough measurements and what the room is for, and the plan comes back in the same thread.

Most of the minute was waiting. The worker checks its queue once a minute, so the message sat until the next check; the decision itself took about two seconds2. A push loop could bring the total down to a few seconds. We have not built one yet: a short delay costs the customer nothing they notice, while the one-minute cycle keeps exactly-once handling simple to audit. For a shop that normally replies the following day, an answer inside a minute already feels different.

The message it would not answer

Six minutes later the customer account wrote that it had paid for a plan and wanted a refund. The agent did not answer the question. It told the customer that a person from the team would reply in the same thread, and stopped. It offered no apology that admits fault, no summary of a policy that might be wrong, and no gesture it had no authority to make.

Two rows from the decision log, day 3
Time (UTC)DecisionTierSentIntent
18:45:09Answer3YesAsking about the first room plan
18:50:07Hand over1HeldRefund request

The owner’s email carried the intent, the customer’s words, the agent’s exact acknowledgement and the details collected along the way, such as a name, an item and a date. Quoting the acknowledgement means the owner never contradicts what the customer was already told.

A bug caught on our account, not a client’s

On day two the worker processed the same batch twice. Seven test messages turned into fourteen decisions, and five handover emails went out twice. The cause was ordinary: a folder permission was wrong, the step that marks an event as done failed without an error, and each cycle read the same queue again. Live, one customer would have been greeted over and over.

Fixing the permission was not enough, since permissions can break again. The fix makes correctness independent of it: before any work, the worker claims the event id in a record that can be written only once. When that claim is already there, the event counts as handled, whatever became of the file. We proved it by injecting one event and running the worker twice; the second run found nothing to do.

Two flaws only real messages showed

The customer was left waiting in silence

In the first version, a handover meant total silence toward the customer, acknowledgement included. Live, the flaw was obvious: the customer asking for a refund was left on read. Acknowledgements for handovers and refusals now go out. Such a message states nothing and commits to nothing; it says a person will answer in the thread. The facts still wait for the owner.

The urgent email was held back

To stop a chatty customer from producing five emails in five minutes, handover emails were limited to one per conversation per hour. A real sequence exposed the problem: a casual message used up the hour, and the refund request six minutes later sent no email. The most urgent tier now always sends; the limit applies to routine tiers only.

Each fix was one line. Neither flaw was visible until real messages arrived.

What it left for a person

Nine conversations were recent on the account at the time. Three got a drafted reply3. The remaining six held only pictures or shared posts, with no text to base an answer on. The agent flagged those for the owner, and the report names each one instead of burying the count.

An agent that replies to pictures it cannot read makes confident mistakes. Telling the owner that six threads are theirs costs less than one wrong answer.

Built, and deliberately not built

Built
  • Replies drawn from the approved list and written as the owner writes
  • Handover by email with the full context
  • A log row for each decision, feeding the reports
  • A rollout that starts paused and runs behind an allow-list
Deliberately not built
  • A client dashboard: owners reply inside Instagram
  • Replies to messages that are only a photo
  • Answers about money, from any agent
  • Any action that cannot be undone, taken without a person

Leaving out a dashboard was deliberate. Asking an owner to learn another inbox hands them extra work. Owners answer inside Instagram, and once the agent sees a human reply in a thread it stays out of that thread.

What a client goes through

  1. Connect in one tap. A private link opens the platform’s official consent screen on the owner’s phone. We never see a password, and the owner can revoke the connection from their own settings at any time.
  2. One written questionnaire. The owner supplies the answer list: the questions customers send, answered in the owner’s words, plus the topics that always go to a person and the topics the agent must never touch. A clinic rules out clinical judgement, a restaurant rules out allergen promises, and every business rules out money.
  3. A silent week, if wanted. The agent drafts everything and sends nothing while the owner reads what it would have said. The switch flips only when the drafts read right, and a personal reply to any thread still silences the agent there.
  4. Reports, daily and monthly. The log that produced every figure here produces the client’s too: replies sent, threads handed over, and questions with no answer yet that belong on the list. Each figure is printed next to its definition.

The steps of the engagement itself are set out in how an engagement runs, and the service on the Inbox Agent page.

Definitions and sources

First reply
Time from the message arriving to the reply being sent, including the one-minute polling cycle.
Handover
A conversation the agent stepped back from and sent on to a person, together with everything said so far.
Held
A reply the engine drafted that the send step refused to transmit.

Figures: 1–2 worker log, day 3, 18:44 to 18:45 UTC; 3 inbox read with skip counters, day 2; 4 the handover list is enforced in the send step. Automated replies are disclosed to customers.

Written by Sathish Balakrishnan, Founder of iHayz

Case studies are written from the system’s own log by the people who built it. The work is contracted through Media Experts LLC in the USA or Media Experts in India. About the company

Start

Not sure what you need? Ask in writing.

Describe the work in a few lines. We will reply in writing within one business day with what we would build and how.