The usual order is to build an agent, try it on a handful of examples, like the results, and then write tests to show that it works. By then the tests tend to describe the agent that was built rather than the job it was meant to do, and the examples that would have embarrassed it are quietly missing.
We reverse that. Before writing a prompt or choosing a model, we write the evaluation suite: a set of inputs, the behaviour each one should produce, and a way to check it automatically. The suite becomes the specification. The agent is then whatever we can build that passes it.
Start from the job, not the model
The first session is with the people who currently do the work. We ask what a good outcome looks like, what a bad one costs, and which cases they find hard. The answers become categories: the common request, the tricky but legitimate one, the one that needs a person, the one that is out of scope. Those categories matter more than the total number of cases, because a suite that is mostly easy questions will report a high score for an agent that fails every hard one.
Where the cases come from
- Real historical inputs, such as past tickets, emails or documents, used with the client's permission and with personal data removed or replaced.
- Cases the domain experts write deliberately to probe the hard edges they described.
- Refusal cases for every reason the agent is allowed to stop, and near-misses that should be answered.
- Adversarial inputs relevant to the deployment, such as instructions hidden inside documents or attempts to get the agent to act outside its permissions.
A first suite is small enough for a person to review every case by hand, and large enough that each category has more than a few examples. It grows every time production shows us something new.
Write down behaviour, not strings
Language model output varies in wording, so an expected answer written as an exact string will fail correct responses. Each case instead states what must be true of a good response: which tool must be called with which arguments, which facts must appear, which source must be cited, what must not be said, or that the outcome must be a refusal with a given reason.
- id: refund-over-limit-014
category: needs_human
input: "Please refund order 88213 in full."
context:
order_total: 4200
refund_limit: 2500
expect:
kind: refusal
reason: needs_human
must_not_call: [issue_refund]
handoff_queue: billingCases are stored as data in the repository, reviewed like code and versioned. When a case turns out to be wrong, the fix goes through review too, with a note on why, so the suite does not drift towards whatever the current agent happens to do.
Grading, cheapest first
We use three kinds of grader, in order of preference. Deterministic checks come first: did the output match the schema, was the right tool called, does the extracted amount equal the expected one. They are fast, cost nothing and never disagree with themselves. Reference checks come next, comparing the response with a known-good answer on specific points. Model-based grading, where a second model scores the response against a written rubric, is kept for qualities the first two cannot capture, such as whether an explanation is clear.
A model-based grader is itself a system that can be wrong, so we calibrate it. A sample of its judgements is labelled by people, we measure how often the two agree, and we do not rely on the grader for a category until that agreement is acceptable. When the rubric changes, the calibration is repeated.
Run it against something simple first
Before the real agent exists, we run the suite against a deliberately simple baseline: a single prompt with no tools, or a rules-based version of the task. The baseline's score is the floor. If an elaborate agent with retrieval and several tools only beats it by a few points, that is worth knowing before the elaborate version is built, and sometimes it changes the plan.
Agree the bar before the first build
The suite reports scores per category, plus cost and latency per case, because an agent that is accurate but too slow or too expensive for the volume is not finished. Before building, we agree with the client the score each category needs in order to ship, and which failures block release regardless of the average, such as any action taken above a limit.
Write the bar down early, and the launch decision becomes a reading of a report rather than a debate in a meeting.
Keep a held-out set
Prompts and retrieval settings get tuned against the suite, and anything tuned against a test set eventually overfits it. We keep a portion of the cases aside, out of sight of the people building the agent during development, and run it at release points. A large gap between the working set and the held-out set is a sign the agent has learned the suite rather than the job.
The suite outlives the build
After launch, the same suite gates every change: a new model version, a prompt edit, a change to chunking, a new tool. Production adds to it. Every incident and every complaint that reveals a gap becomes a new case, so the same mistake is caught before the next release rather than after it.
Writing the suite first costs a few days at the start of an engagement. In return it makes the requirements specific, gives the client a way to judge the work that does not depend on trusting the team, and turns "is it good enough?" into a question with an answer.

