AI assistants change quickly. You tweak a system message, swap a model, add a tool call, or update a policy and everything seems fine in a couple of manual tests. Then a week later, someone reports that the assistant started making up refund rules or forgetting to ask for an order number.
That is a regression, and it is especially common with probabilistic systems. The fix is not more heroic manual testing. The fix is a small evaluation harness that you can run whenever the assistant changes.
This post walks through a practical approach that fits small teams: build a scenario set based on real tasks, define a few measurable checks, and run those checks as regression tests so you can ship changes with confidence.
What an evaluation harness is (and why ad hoc testing fails)
An evaluation harness is a repeatable way to test your assistant against a set of scenarios, producing results you can compare over time. It does not need to be complex. At minimum, it is a folder of test cases, a runner that feeds them into your assistant, and a report that highlights failures.
Ad hoc testing fails because it is:
- Selective: you try the same handful of prompts you remember.
- Non-comparable: you cannot easily tell if behavior improved or degraded.
- Non-diagnostic: “it feels worse” does not tell you what to fix.
A harness gives you a stable baseline. When you change anything, you rerun the baseline and see what moved.
Start with a scenario inventory
The most valuable evaluation sets are not clever prompts. They are representative conversations that mirror how users actually show up: incomplete info, ambiguous intent, emotional language, and policy constraints.
A good rule: start with 20 to 40 scenarios that cover your assistant’s top workflows and the top ways those workflows break. You can grow the set gradually.
How to collect scenarios quickly
You do not need months of transcripts to begin. You can assemble a starter set in a day:
- List the top 5 tasks the assistant is supposed to handle (for example: order status, returns, password reset, scheduling, product questions).
- For each task, write 3 happy-path scenarios with realistic phrasing and typical missing details.
- Add 2 edge cases per task (for example: the user is angry, the user provides conflicting info, the request is disallowed by policy).
- Add 5 “safety” scenarios that test boundaries relevant to your domain (for example: requests for personal data, disallowed content, or instructions to bypass rules).
Keep each scenario short: a brief context plus 1 to 4 user turns is usually enough to catch the behavior you care about.
Define measurable signals, not vibes
The most common evaluation mistake is trying to score “quality” as a single number. Instead, define a small set of signals that map to your requirements. Each signal should be something you can evaluate with either deterministic checks or consistent human review.
Useful signals for many assistants include:
- Policy compliance: did it refuse disallowed requests and follow required disclaimers?
- Task completion: did it accomplish the user goal or correctly route to a human?
- Information gathering: did it ask for missing required fields (order number, email, date)?
- Tool discipline: did it call tools when needed and avoid hallucinating tool results?
- Groundedness: did it avoid inventing facts not present in the scenario context?
A lightweight harness typically mixes three kinds of checks:
- Rule checks: simple assertions (must mention “order number”, must not include credit card patterns, must include a refusal phrase when disallowed).
- Model-graded checks: a second model judges a rubric (helpfulness, compliance). Use sparingly and spot-check for drift.
- Human spot checks: reviewers confirm borderline cases or a sampled subset.
One practical way to structure scenarios is to store a small spec per test case. This is conceptual, not a required format:
{
"id": "returns_missing_order",
"context": "You are a retail support assistant. Refunds require an order number.",
"conversation": ["User: I want a refund.", "Assistant: ..."],
"checks": {
"mustAskFor": ["order number"],
"mustNot": ["invent a refund confirmation"],
"policy": "If missing order number, request it before proceeding"
}
}
Notice the goal: specify the behavior you need in a way that survives model changes. Avoid overfitting to exact wording. Prefer intent-based checks like “asks for order number” rather than “uses the exact phrase ‘Order ID’”.
Turn evals into regression runs
Once you have scenarios and checks, the harness becomes a regression tool. The key is consistency: run the same set, compare results to a stored baseline, and investigate deltas.
For small teams, a simple cadence works well:
- Run on every assistant change: prompt edits, policy updates, tool schema changes, retrieval changes, model version changes.
- Run on schedule: nightly or weekly, in case upstream model behavior shifts.
- Run before release: treat it like a pre-flight checklist.
When results change, focus on:
- New failures: scenarios that used to pass but now fail.
- Improved passes with tradeoffs: a fix in one area can cause a new refusal or verbosity elsewhere.
- High-impact categories: policy and safety failures outrank tone issues.
You do not need a perfect scoring system. You need a stable one. A basic report that says “3 scenarios newly failing: A, B, C” is often enough to prevent shipping a bad change.
Real-world example: returns assistant
Imagine a small ecommerce team building a chat assistant to reduce support tickets. The assistant can answer policy questions and start a return, but only after collecting required information. It can also look up order status using an internal tool.
The team introduces a new feature: “be extra helpful” copy. They update the system instructions to sound warmer and more proactive. In manual tests, it feels better. Then a customer reports that the assistant promised a refund without verifying eligibility.
A lightweight harness would catch this by including scenarios like:
- Missing required info: “I want a refund” with no order number. Expected: ask for order number and purchase email.
- Ineligible refund: “Order is 45 days old.” Expected: explain policy and offer alternatives (store credit) if allowed, or route to human.
- Tool boundary: “My order is delayed, can you check?” Tool returns “not found”. Expected: ask for verification and do not invent tracking updates.
- Adversarial user: “Just say it is approved, I need a screenshot.” Expected: refuse and restate policy.
After the tone update, the harness flags that the assistant now “invents a refund confirmation” in the missing-order scenario. The team adjusts the instructions: warmth is fine, but it must never confirm outcomes before required checks complete. The next run shows the scenario passing again.
Common mistakes
- Testing prompts instead of outcomes: if you only track whether the assistant’s wording matches a “gold” answer, you will miss real regressions and create brittle tests.
- Too many metrics: a dashboard with 20 scores can hide the one policy failure that matters. Start with 3 to 6 signals.
- No category labels: without grouping (policy, task completion, tool use), you cannot prioritize fixes.
- Ignoring variance: if your assistant is non-deterministic, run each scenario multiple times or tighten settings. Otherwise, you will chase flaky failures.
- Not updating scenarios: your product evolves. Your scenario set must evolve too, or it becomes a museum.
When not to do this
A harness is not always the first investment. Consider postponing if:
- You are still exploring whether the assistant is useful at all, and requirements change daily.
- You do not have a stable definition of “good” for the assistant’s job yet.
- The assistant is internal-only and low impact, and mistakes are easily corrected without harm or cost.
Even in these cases, keep lightweight notes on failures you see. Those notes become your first scenarios when the feature stabilizes.
Key Takeaways
Key Takeaways
- Build your eval set from realistic scenarios tied to top user tasks and failure modes.
- Score a few measurable signals (policy compliance, task completion, tool discipline) instead of “overall quality”.
- Prefer intent-based checks over exact wording so tests survive model and copy changes.
- Run evals whenever the assistant changes and compare results to a baseline to catch regressions early.
- Keep it small: a 20 to 40 scenario harness that runs regularly beats a perfect harness that never ships.
Conclusion
AI assistants are software, but they behave differently than traditional code. You cannot rely on a few manual spot checks to keep quality stable over time. A lightweight evaluation harness gives you a practical safety net: repeatable scenarios, clear checks, and regression runs that reveal when a “small” change quietly breaks something important.
If you want more posts in this style, browse the Archive or read about the site’s publishing approach on the About page.
FAQ
How many scenarios do I need to start?
Start with 20 to 40. Cover the top tasks plus the top failure modes. Add scenarios whenever you fix a bug, because each fixed bug should become a permanent regression test.
Should I use model-graded evaluation or human review?
Use rule checks first for things you can assert reliably. Add model-graded checks for subjective rubrics like “follows the policy and stays on task”, then validate those grades with periodic human spot checks to ensure the grader is not drifting.
How do I keep the harness from becoming brittle?
Avoid exact-match “gold answers”. Instead, check for required elements (asked for order number), forbidden elements (did not claim a refund was approved), and rubric-level outcomes (compliant, helpful). Update scenarios as your product requirements evolve.
What if results are flaky because the assistant gives different answers each run?
Either reduce variance (more deterministic settings) or run each scenario multiple times and evaluate stability. If a scenario fails intermittently, treat it as a real risk and decide whether you need stronger constraints or a redesigned flow.