“It looks better in my testing” is not an evaluation strategy, and it’s exactly how prompt and model changes quietly regress production quality. LLM evaluation is the discipline of turning “seems good” into a number you can track across changes — the same instinct that made unit tests non-negotiable for regular code, applied to a component whose output is nondeterministic by design.

Why “It Seems to Work” Isn’t an Eval

Manually reading transcripts catches obvious failures and misses everything subtle — a retrieval system that answers 90% of common questions well and silently fails on a long tail you never happened to try. Worse, it doesn’t survive a change: without a fixed, repeatable set of inputs and a scoring method, there’s no way to say whether a new prompt or a new model version made things better or worse, only whether it feels different on the handful of examples you happened to glance at.

Building an Eval Set

A good eval set is built from real failure cases, not invented ones. Start by logging production traffic (with appropriate privacy handling), and pull the cases where users complained, retried their query, or gave negative feedback — those are exactly the inputs your current system handles badly, which makes them the highest-value additions to the eval set. Supplement with synthetic edge cases for scenarios you know matter but haven’t seen enough real examples of yet.

eval_cases = [
    {
        "id": "billing-double-charge-001",
        "input": "My invoice shows twice the usual amount, what happened?",
        "expected_category": "billing",
        "expected_facts": ["reset window", "60-second"],
    },
    {
        "id": "ambiguous-002",
        "input": "it's not working",
        "expected_category": "bug",
        "expected_behavior": "ask_clarifying_question",
    },
]

Deterministic Checks vs Model Graders

Not every output needs a model to grade it. Structural properties — is the output valid JSON, does it match a required schema, is a required field present, does a numeric answer fall within a tolerance — should be checked with plain code. It’s faster, free, perfectly reproducible, and has zero risk of the grader itself being wrong in a way that’s hard to detect.

Check type Use for Cost
Exact match / regex Structured fields, classification labels Free, instant
Schema validation JSON output shape Free, instant
Deterministic function Numeric tolerance, business rule checks Free, instant
Embedding similarity Loose semantic match to a reference answer Cheap, fast
LLM-as-judge Open-ended quality, faithfulness, tone Slower, costs tokens, needs its own validation

Reach for a model grader only once deterministic checks run out — for genuinely open-ended qualities like “is this answer faithful to the retrieved context” or “does this response match our support tone,” where there’s no fixed string to match against.

LLM-as-Judge

Using a model to grade another model’s output works, but it inherits the same reliability concerns as any other LLM call — it needs a clear rubric, and it needs its own accuracy checked against human judgment before you trust it. A vague grading prompt (“is this a good answer?”) produces inconsistent scores; a specific rubric with concrete pass/fail criteria produces something closer to a reliable metric.

Promptclaude-sonnet-5
You are grading whether an AI support response is faithful to its source context.
Respond with JSON: {"faithful": true|false, "reason": "..."}

Source context:
"""
Refunds are processed within 5-7 business days after approval.
"""

AI response: "Your refund will be processed within 5-7 business days once approved."

Output

          {"faithful": true, "reason": "The response accurately restates the timeline and condition from the source context without adding unsupported claims."}
        

Regression Testing Prompts

Treat prompt and model changes the way you’d treat a code change: run the full eval suite before merging, not after noticing a problem in production. A prompt change that improves the three examples you tested it against can regress a dozen cases you didn’t think to check — that’s exactly what a fixed eval set catches and ad hoc testing doesn’t.

# Run the eval suite against a candidate prompt version, diff against baseline
python run_evals.py --prompt-version candidate --baseline main --output results.json

Track scores over time, per category, not just as one aggregate number. An aggregate score can stay flat while a specific category — say, ambiguous queries — quietly gets worse, masked by improvement elsewhere in the set.

Statistical noise is easy to underestimate here. Sampling temperature and inherent model variance mean the same prompt against the same eval case can score differently across runs, so a single run’s pass rate is a noisy estimate, not a fact. Run the suite multiple times per candidate, or run at low sampling variance for the eval specifically, before treating a small score delta between two prompt versions as a real improvement rather than noise.

Evaluating Agents and Multi-Step Systems

Agents need evaluation at two levels: the final outcome (did the task actually get completed correctly) and the trajectory (were the intermediate steps reasonable, or did the agent get the right answer through an unreliable path that will fail differently next time). A trajectory eval checks things like: did it call the right tools, in a sensible order, without redundant calls, and did it recover reasonably when a tool returned an error.

Takeaway

Evaluation is what turns prompt and model changes from a guess into a measured decision. Build the eval set from real failures, use deterministic checks wherever the output has a checkable structure, reserve model graders for genuinely open-ended judgments, and validate those graders against human labels before trusting the numbers they produce.