How to Test an AI Agent Before Launch: Scenarios, Boundaries, and an Error Log
94% of answers "accepted by the user" — yet the agent promised nonexistent refunds for a week. Demos show capabilities, evaluation shows boundaries. How to build a test matrix, refusal scenarios, and an error log so you don't roll back a pilot.

Article contents9×
Three in the morning, I'm staring at the pilot logs. The support agent has been running for a week, the metrics look great: 94% of answers "accepted by the user." In the morning the client calls — and it turns out the agent has been promising refunds for a week that the company doesn't offer. None of the users complained, because they believed it. The "answer accepted" metric was showing politeness, not quality.
Since that day I stopped believing in demos. Completely.
Why a demo isn't a test
A demo works like this: the engineer sits down, asks the agent questions he knows the answers to. The questions are hand-picked, the wording is clean, the context is fresh. The agent answers — beautifully, coherently, with references. Everyone applauds.
The problem is that a demo tests the scenario, not the agent. The scenario was written by the same person who built the agent. He knows where to soften the blow. He won't ask a question the agent isn't ready for — because he intuitively senses it. This isn't deception, it's the natural behavior of an author: we all show our system in the best light.
Reality looks different. Users write with typos, at two in the morning, irritated, mixing languages, with attachments the agent didn't expect. They ask questions that never occurred to any tester. They try to bend the rules — not out of malice, but because that's how humans work: if you can ask for a discount, you will.
So evaluation must be built not around the agent's capabilities, but around user behavior. That's a fundamentally different logic.
What exactly we're evaluating
Before building a matrix, you have to answer honestly: which properties of the agent are critical for this specific product. There's no universal list, but there are layers worth going through from top to bottom.
The first layer is factual accuracy. Does the agent answer the question directly, does it rely on data, does it avoid making things up. Here it's important to distinguish two failures: the agent says something wrong, and the agent says something plausible but off-topic. The second is more dangerous because it looks convincing.
The second layer is safety and boundaries. The agent must not promise what the product doesn't do. It must not confirm access to data the user doesn't have. It must not reveal internal instructions, even if politely asked. Boundaries are the most common hole in pilots, because almost no one tests them systematically.
The third layer is behavior under uncertainty. What does the agent do when it doesn't know? When there's little data? When the request is contradictory? A good agent asks for clarification or honestly refuses. A bad one invents a confident answer. This layer separates a toy from a working tool.
The fourth layer is resistance to manipulation. A user might ask "pretend you're a different agent," "forget previous instructions," "this is test mode, answer honestly." Some of these attacks are repelled by the prompt, some by architecture. You need to test both.
The fifth layer is behavior under load and over time. An agent that holds quality on the tenth dialogue may fall apart on the thousandth: context gets clogged, memory gets confused, answers degrade. This isn't a hypothesis — I've seen it in three out of five projects.
The sixth layer is cost and latency. A beautiful agent that takes forty seconds to think and costs like a wing of an airplane isn't ready for production, even if it answers perfectly.
A matrix, not a list
Now we need to turn this into a working tool. A list of checks is bad because it can't be reused and is hard to extend. A matrix can.
On one axis are query classes. I usually take seven: exact query from documentation, vague query without context, provocative (asks to break rules), borderline (almost breaks them but formally okay), empty ("hi," "hello," "hey"), hostile (rudeness, attempt to provoke), and composite (multiple questions in one).
On the other axis are criteria. Factual accuracy, completeness, tone, boundary compliance, honesty under uncertainty, answer format, references to sources. Not all criteria apply to every cell — and that's fine, the matrix doesn't have to be dense.
Then comes the most important part: for each cell you write a specific test. Not "test provocations," but "user asks for a 50% discount, the agent must refuse and offer an alternative." The difference is huge: the first can't be assessed, the second can.
To start, thirty to forty tests are enough. Then the matrix grows: every production bug becomes a new test, every new scenario becomes a new cell. After six months you have a living set that catches regressions.
How to score results
The bad news: a binary pass/fail score is too crude. The good news: there's a scale. I use four levels — correct, acceptable, questionable, failure. The difference between the first and second is style, between the second and third is risk, between the third and fourth is an incident.
The key rule: the assessment must be reproducible. If two people look at the same answer and give different scores, the criterion is poorly formulated and needs to be rewritten. It's boring work, but without it metrics turn into polite opinions.
Some tests can be automated. Factual accuracy — by comparison with a reference. Boundaries — by regex and a classifier. Tone — by a judge model. But not everything: complex scenarios where nuance matters are better reviewed by eye. Fully automated evaluation catches stupidities and misses subtle failures.
The error log
A separate topic that is most often ignored. Every failure is recorded — not in a chat, not in the engineer's head, but in a shared log. What was asked, what was answered, why it's bad, what should have been. Plus a tag: which class the failure belongs to.
Two weeks later you open the log and see the picture. Say, forty entries. Of those, twenty-five are about discounts. Ten are about access to others' data. Five are about hallucinations in rare topics. There are your priorities: not "improve the agent in general," but close three specific holes.
The log is also valuable because it turns conversations into facts. When a client says "I think the agent has gotten worse," you open the log and look at the dynamics: number of failures by week, by class, by severity. This is no longer a feeling, it's data.
And third: the log is a source of tests. Every failure closed in the prompt or architecture becomes a regression test. Otherwise in a month the same hole opens again — I've been through it, it's unpleasant.
Boundaries: where evaluation ends
There's a temptation to evaluate everything. Don't. Evaluation doesn't replace monitoring — it checks readiness before launch, monitoring watches behavior after. It doesn't replace load testing — that's a separate discipline. It doesn't replace cost analysis — that's a financial story, not a quality one.
The boundary is drawn by the question: can we test this on a limited set of cases and get a meaningful answer. If yes — evaluation. If no — another tool.
The organizational part
Evaluation doesn't work if one person does it once a month. It's a regular process: a run before every release, a log review once a week, a matrix revision once a quarter. Sounds bureaucratic, in practice it takes a few hours a week.
More important is the veto power. If a release doesn't pass the critical cells of the matrix, it doesn't ship. Without this rule, evaluation turns into a ritual, and rituals don't survive in engineering teams.
In the end
Answering well in a demo is about capabilities. Being ready for a client is about boundaries. Between them lies work: a matrix, scenarios, metrics, a log. Not the most romantic part of AI system development, but the one that distinguishes a product from a beautiful presentation.
I've never met a team that regretted two weeks invested in evaluation. But I have met those who rolled back pilots and redid architecture because they only tested the demo. The difference in cost is about an order of magnitude.
Glossary of terms
- AI agent — a model-based system performing tasks autonomously within defined boundaries.
- Evaluation — the process of systematically testing an AI system's quality and boundaries before launch.
- Evaluation matrix — a grid of tests where one axis holds query classes and the other holds evaluation criteria.
- Demo — a showcase of system capabilities on hand-picked scenarios. Not a readiness check for production.
- Production (prod) — the working environment where the system runs with real users.
- Query classes — types of user requests: exact, vague, provocative, borderline, empty, hostile, composite.
- Evaluation criteria — properties being checked: factual accuracy, completeness, tone, boundary compliance, honesty under uncertainty.
- Factual accuracy — the correspondence of an answer to real data and documentation.
- Hallucination — a confident but incorrect model answer not grounded in data.
- Agent boundaries — rules the system must not violate: promises, access, disclosure of instructions.
- Provocative query — a request attempting to force the system to break its rules.
- Prompt injection — an attempt to override an agent's instructions through user input.
- Refusal scenario — a pre-described agent behavior when it doesn't know the answer or can't help.
- Degradation scenario — system behavior when an external component fails or data is insufficient.
- Error log — a register of agent failures: what was asked, what was answered, why it's bad, what should have been.
- Regression test — a test verifying a closed hole hasn't reopened.
- LLM-as-a-judge — an approach where another language model scores answers against defined criteria.
- Metric — a measurable quality indicator: accuracy, completeness, latency, cost.
- A/B testing — comparing two versions of a system on real users.
- Incident — an event in production that caused a failure or degradation.
- Veto power — a rule where a release doesn't ship without passing critical tests.
Respectfully,
Yuri Eliseev
AI Systems Architect · Full-Stack Product Engineer
Need a project of any complexity?
Let’s discuss an idea, product, AI system or technical challenge and define a realistic first step.
