How to test an LLM feature before it ships

Code on a monitor viewed through glasses

A practical guide to LLM evaluation: golden datasets, automated metrics vs human review, prompt regression tests, guardrails and cost and latency budgets.

Bhaskar Bhatt8 min read

LLM evaluation is how you find out, before customers do, whether a feature built on a language model gives answers that are correct, safe, fast enough and affordable. In practice it means a fixed set of real test cases, a mix of automated checks and human review, and a rule that no prompt or model change ships without passing that set.

On this page
  1. What does LLM evaluation actually involve?
  2. How do you build a golden dataset that’s worth trusting?
  3. Automated metrics or human review: which should you rely on?
  4. How do you run regression tests on prompts and models?
  5. Which guardrails should be tested before launch?
  6. How should you set cost and latency budgets?
  7. What does “good enough to ship” mean for an LLM feature?
  8. When isn’t a full evaluation setup the right approach?
  9. Frequently asked questions

Traditional tests assume the same input gives the same output. LLM features don’t. The same question can produce two different answers, a provider can update a model under the same name, and a one-word prompt edit can fix one case while breaking ten others. A handful of unit tests and a demo won’t catch that.

What does LLM evaluation actually involve?

It involves three things: a set of test inputs that represent real use, a definition of what a good output looks like for each one, and a repeatable way to score outputs against that definition. Everything else, from dashboards to tooling choices, sits on top of those three. If one is missing, the scores don’t mean much.

Most teams pick an evaluation tool first and only then ask what they’re measuring. Do it the other way round. Write down the job the feature does in one sentence (“answer billing questions using our help centre, and hand off anything about refunds”), then list the ways it could fail that job. Wrong facts, a made-up policy, a missed handoff, an answer in the wrong format, a leaked internal note. Each failure mode becomes something you test for.

How do you build a golden dataset that’s worth trusting?

Collect real inputs, not invented ones. Pull them from support tickets, search logs, sales emails or a pilot group, remove personal data, and have someone who knows the domain write or approve the expected answer for each. Start with 50 to 200 cases covering common requests, known hard cases and things the feature must refuse.

A golden dataset is only as good as the people who labelled it. Some practical rules we follow:

  • Use real phrasing. Customers write short, vague, misspelt questions. Test cases written by engineers are usually too clean.
  • Tag every case. Mark the category, difficulty and expected behaviour (answer, ask a clarifying question, refuse, hand off). Tags let you see that the feature is fine on billing but weak on account access.
  • Write expected facts, not expected wording. For open-ended answers, list the facts that must appear and the claims that must not. Matching exact sentences will fail good answers.
  • Keep a held-out slice. If the team tunes prompts against every case, the set stops telling you how the feature handles questions it hasn’t seen.
  • Version it. Store the dataset in the repository or a versioned bucket, with a changelog. A score is meaningless if nobody knows which version of the set produced it.

Expect the set to grow. Every production bug that gets past it should become a new test case the same week.

Automated metrics or human review: which should you rely on?

Both, for different jobs. Automated checks are cheap and run on every change, so use them for anything you can define precisely. Human review is slow but catches tone, nuance and subtle errors. Use people to calibrate the automated checks and to review a sample before each release, not to score every run.

MethodGood forWatch out for
Deterministic checks (schema, regex, exact match)Format, required fields, classification labels, forbidden strings, citations presentSays nothing about whether the content is right
Reference comparisonExtraction and short answers with one correct valuePenalises correct answers phrased differently
LLM-as-judge with a rubricGroundedness, completeness, policy compliance on open answers, at scaleJudge bias and drift; needs checking against human scores
Human reviewTone, domain correctness, edge cases, final release sign-offSlow, inconsistent between reviewers without a clear rubric
Live signals (thumbs, edits, escalations)Real-world quality after launchSparse and skewed towards unhappy users

Using one model to grade another is practical, but don’t treat the grade as truth. Research on LLM judges, including the MT-Bench and Chatbot Arena study by Zheng et al., documents position, verbosity and self-enhancement bias. Give the judge a narrow rubric with yes/no questions (“Does the answer state anything not supported by the retrieved documents?”), and check its scores against human ratings on a sample before you rely on it. If the judge and your reviewers disagree often, fix the rubric before you trust the numbers.

How do you run regression tests on prompts and models?

Treat prompts, model versions, retrieval settings and tool definitions as code. Store them in version control, and run the full evaluation set automatically whenever any of them changes. Compare the new scores with the last approved baseline, per category, and block the merge if a key metric drops past an agreed threshold.

The steps we put into a CI pipeline for an LLM feature look like this:

  1. Pin everything. Use dated model versions where the provider offers them, and record temperature, system prompt and retrieval parameters with each run.
  2. Run the set more than once. Outputs vary between runs, so score each case several times on key metrics and look at the spread, not a single number.
  3. Compare against the baseline. Report changes per tag. An average that holds steady can hide a large drop in one category.
  4. Diff the failures. Show the reviewer which cases flipped from pass to fail, side by side with the old output. That is where the useful information is.
  5. Gate the release. Agree in advance which metrics are blocking and which are advisory, and write the thresholds down.
  6. Re-run on a schedule. Hosted models can change behaviour without a code change on your side. A nightly or weekly run against production settings catches that.

If your team already has a delivery pipeline, this slots in as another test stage. Our MLOps and AI infrastructure work usually starts by adding exactly this stage, plus tracing so failures can be replayed.

Which guardrails should be tested before launch?

Test the guardrails the same way you test answers: with cases designed to break them. That means prompt-injection attempts in user input and retrieved documents, requests for data the user shouldn’t see, off-topic or harmful requests, and outputs that downstream code will parse or execute. Each guardrail needs passing and failing examples in the evaluation set.

The OWASP Top 10 for LLM Applications is a sensible checklist to work from. It covers prompt injection, sensitive information disclosure, improper output handling, excessive agency and unbounded consumption, among others. For each item that applies, write a few adversarial test cases and keep them in the regression set permanently. Untested guardrails tend to vanish in prompt rewrites.

Two things are easy to miss. First, test the refusal rate as well as the leak rate. A feature that refuses half of legitimate requests is also broken. Second, test what happens when the model provider is slow or down. The fallback path is part of the feature. If the feature touches personal data or regulated decisions, our AI governance and security team can map these tests to the controls your auditors will ask about.

How should you set cost and latency budgets?

Set them before you tune quality, because they constrain every other choice. Decide the slowest acceptable response time for the user experience, usually as a 95th-percentile target, and the most you can spend per request or per completed task. Then measure both on every evaluation run, alongside accuracy.

Cost and latency move with decisions that look like quality decisions: a larger model, longer prompts, more retrieved chunks, extra verification calls. A change that adds a little accuracy but doubles response time may not be worth shipping. Record tokens, model calls and wall-clock time per case in the same report as the quality scores, so the team argues about real numbers.

Also cap the worst case. Put limits on output tokens, retries and tool-call loops, so one strange input can’t run up a large bill or hang a request.

What does “good enough to ship” mean for an LLM feature?

It means the feature beats the current alternative on the cases that matter, fails safely on the rest, and stays inside its cost and latency budgets. “Good enough” is a written threshold agreed with the business owner before testing starts, not a feeling after a demo. No LLM feature reaches a perfect score.

We usually write release criteria in three tiers. Hard gates must pass every time: no leaks of other users’ data, no unsafe actions, correct handoff on regulated topics. Quality targets are scores on the main categories, compared with the current process, whether that’s a human agent, a search box or nothing at all. Budgets cover cost and latency. A feature can launch with a quality gap in a low-risk category if the team has a plan to close it. It shouldn’t launch with a hard gate failing.

Be honest about the comparison. If the feature is slightly worse than a support agent but answers instantly at 2 a.m., that may be the right trade for a limited rollout with a human fallback. If it’s worse and slower, it isn’t ready. For how we scope this kind of feature from the start, see our GenAI solutions page and the post on building a GenAI MVP in 4–8 weeks.

When isn’t a full evaluation setup the right approach?

When you’re still exploring whether the idea works at all. In the first week of a prototype, reading 20 outputs by hand is faster and more informative than building a pipeline. Formal evaluation starts paying for itself once you’re choosing between prompts, models or retrieval designs, or once real users will see the output.

It’s also overkill for internal tools where every output is reviewed by a person before it goes anywhere, such as a drafting assistant for your own team. There, a smaller sample review each month and a feedback button may be enough. And if you can’t get domain experts to agree on what a correct answer looks like, fix that first. No metric will rescue a task nobody can define. For agents that take actions, rather than features that only return text, the stakes are higher; our post on taking AI agents from PoC to production covers the extra controls.

Frequently asked questions

How many test cases does an evaluation set need?

Enough to cover each category and failure mode with several examples. For a narrow feature, 50 to 200 well-chosen cases is a workable start. Coverage of the hard cases matters more than the total, and the set should grow every time a real failure slips through.

Can we use synthetic test data generated by an LLM?

Yes, as a supplement. Generated cases are useful for adding variations and adversarial inputs quickly. They tend to be cleaner and more predictable than real user input, so keep real examples as the core of the set and have a person review generated ones before adding them.

Should we evaluate the retrieval step separately from the answer?

Yes, for any feature built on retrieval-augmented generation. Check whether the right documents were retrieved before judging the answer. If retrieval misses, no prompt change will fix the output, and measuring the two together hides where the problem is.

Scroll to Top