LLM Evaluation: 7 Critical Methods That Actually Work

LLM evaluation means scoring model outputs against a gold set. I built a Python harness with exact-match, token-F1 and keyword scoring, ran 8 cases, and show why the metrics disagree and which to trust.
LLM evaluation: same answers score 75% or 28% depending on the metric

Published: 2026-10-03

LLM evaluation is how you measure whether a model’s output is actually correct, rather than just fast or cheap. It is the half of quality that monitoring cannot see. I built a small evaluation script in Python, scored eight model answers against a gold set with three different methods, and watched them disagree: exact-match said 75% correct, token-F1 averaged a bleak 0.28, and the harness still cleanly caught both wrong answers. This guide is that hands-on walk-through, including why the scores diverge and which number you should trust for which job.

Evaluation is the companion to LLM observability: observability tells you a call was fast and cheap, evaluation tells you the answer was right. If you are running a model from my local LLM guide, you need both before you ship anything.

Table of contents

  1. What LLM evaluation actually means
  2. The gold set is the whole game
  3. Three scoring methods I ran
  4. Why the three scores disagreed
  5. When to use LLM-as-judge instead
  6. Building evaluation into your pipeline
  7. Evaluation vs observability
  8. FAQ

What LLM evaluation actually means

LLM evaluation is the process of comparing model outputs against expected results to produce a score you can track over time. The point is regression detection: when you change a prompt, swap a model, or bump a version, a good evaluation suite tells you immediately whether quality went up or down, instead of you discovering it from user complaints a week later.

The naive approach is to eyeball a few outputs and declare victory. That falls apart the moment you have more than a handful of cases or more than one person making changes. Real LLM evaluation is automated and repeatable, run on every change, exactly like a unit test suite. Frameworks such as OpenAI’s Evals formalise this pattern.


The gold set is the whole game

Every evaluation needs a gold set: a list of inputs paired with known-good answers. My gold set used eight cases spanning facts and arithmetic as well as code, with the model answers deliberately mixed between correct, verbose, plus a couple that were flat wrong so the scorers had something to separate.

A gold set does not need to be large to be useful. Twenty to fifty well-chosen cases that cover your real traffic beat a thousand random ones. Coverage of the inputs you actually see matters far more than raw volume, because a focused set that mirrors production catches the failures that real users would hit while still running fast enough to execute on every single change. The discipline is to add every production failure you find back into the gold set, so the same mistake can never silently return. That is how an evaluation suite compounds in value.


Three scoring methods I ran

I scored the same eight answers three ways. Each method asks a different question.

LLM evaluation exact match vs token F1
The three methods disagree on the same 8 answers: 75% vs 28%.

Exact match. Does the gold answer appear in the output. Strict and binary, best for facts and short answers. My run scored 75%, catching both wrong answers (the model said Saturn is the largest planet and Sydney is Australia’s capital).

Token F1. The overlap between the gold tokens and the answer tokens, as a harmonic mean of precision and recall. It gives partial credit, which sounds better but has a catch I will come to. Mean F1 across the run was 0.28.

Keyword containment. Does the output contain any significant word from the gold answer. The loosest check, useful as a cheap smoke test. It also scored 75% here.

Here is the per-case picture:

Question Exact match Token F1
Capital of Japan PASS 0.29
2+2 PASS 0.67
Author of 1984 PASS 0.50
Largest planet FAIL 0.00
HTTP not found status PASS 0.29
Python list reverse PASS 0.22
Capital of Australia FAIL 0.00
Boiling point of water PASS 0.29

Why the three scores disagreed

This is the lesson that matters. Exact-match said 75% correct, but mean token-F1 was only 0.28, a number that looks like failure. Both are right, because they measure different things.

LLM evaluation per-case F1 scores
Per case: F1 is low even for correct answers; the two reds are genuinely wrong.

Token F1 is dragged down by verbosity. The answer “The capital of Japan is Tokyo” is completely correct, but it contains six tokens where the gold answer has one, so precision is low and F1 lands at 0.29. The model was right and the metric punished it for being polite. That is the trap: a low F1 does not mean a wrong answer, it often means a wordy one.

Exact match avoided that trap here because it only checks whether the gold string is present. For short factual answers, exact match or keyword containment is the honest metric and F1 is misleading. For longer free-form answers where wording genuinely varies, both F1 and semantic-similarity scoring earn their place. Choosing the wrong metric for the task is the most common evaluation mistake I see.


When to use LLM-as-judge instead

String-based metrics break down for open-ended output like summaries, explanations, or code with many valid forms. For those, the current standard is LLM-as-judge: a second model scores the first model’s output against a rubric. It handles paraphrase and partial correctness that exact-match and F1 cannot.

LLM evaluation metric selection guide
Match the metric to the answer type.

The honest caveats are real. An LLM judge is slower and costs money per evaluation, it can be biased toward verbose or confident answers, and it needs its own validation against human ratings before you trust it. Use it where string metrics genuinely fail, keep deterministic metrics everywhere else, and never treat a judge score as ground truth without spot-checking it against people.


Building evaluation into your pipeline

The payoff comes from running the suite automatically. Wire the gold set into a script that runs on every prompt or model change, the same way tests run on every commit. Fail the change if the aggregate score drops below a threshold you set from the current baseline.

Keep the suite deterministic where you can, because a flaky evaluation run is worse than none. Store each run’s scores so you can see the trend, not just today’s number. And treat the gold set as living: every real-world failure becomes a new case, so your evaluation gets stricter exactly where the model is weak. In practice this means your worst production incident this month becomes next month’s regression test, and the model can never ship that same mistake again without the suite catching it first. Over a few months that turns a thin starter set into a sharp, battle-tested one.


Evaluation vs observability

The two are often confused. Evaluation measures whether the output is correct, offline, against a known answer. Observability measures how a call behaved in production, including its latency and token use, its cost, plus whether it errored. One is about quality, the other about operational health. A model can score perfectly on your evaluation suite and still be too slow or too expensive in production, which is why you track both. I cover the operational half in the LLM observability guide.

Get the next hands-on breakdown

New developer deep-dives on AI, cloud, security and careers — the stuff I actually test. No fluff, unsubscribe anytime.


FAQ

What is LLM evaluation?

LLM evaluation is the practice of scoring a model’s outputs against known-good answers to produce a trackable quality number. Its main purpose is regression detection: knowing immediately whether a prompt change, model swap, or version bump made quality better or worse, rather than finding out from users.

What metrics are used to evaluate an LLM?

Common ones include exact match (is the gold answer present), token F1 (word overlap with partial credit), keyword containment (a cheap smoke test), and LLM-as-judge (a second model scoring against a rubric). Short factual tasks suit exact match; open-ended output needs a judge or semantic similarity.

Why did my F1 score look bad for correct answers?

Token F1 penalises verbosity. A correct but wordy answer such as “The capital of Japan is Tokyo” has many more tokens than the one-word gold answer, which lowers precision and drags F1 down. In my run, correct answers scored as low as 0.22 on F1 while exact match marked them as passes. For short answers, trust exact match over F1.

What is LLM-as-judge?

LLM-as-judge means using a second language model to grade the first model’s output against a rubric. It handles paraphrase and partial correctness that string metrics miss, but it is slower, costs money per evaluation, can be biased toward verbose answers, and must be validated against human ratings before you rely on it.

How big does my evaluation gold set need to be?

Smaller than most people expect. Twenty to fifty cases that mirror your real traffic are more useful than a thousand random ones. The key habit is to add every production failure back into the set, so the same mistake cannot silently return and the suite gets stronger over time.

Is evaluation the same as observability?

They are separate. Evaluation measures output correctness offline against known answers. Observability measures operational behaviour in production such as latency and cost alongside error rates. You need both, because a model can be accurate yet too slow or expensive to ship.

Shares:
Post a Comment

Leave a Reply

Your email address will not be published. Required fields are marked *