LLM evaluation in CI means running a fixed set of realistic questions against your application on every change and failing the build when quality drops. The fixed set is called a golden set. This guide shows how to build one, which metrics to choose, how to gate releases with DeepEval, and how to use Langfuse traces from production to keep the set honest.

Why evaluation belongs in CI

Ordinary code changes when someone edits code. An LLM application changes when a prompt is reworded, a model is swapped or retired, a chunk size is adjusted or a tool description is edited. None of those changes looks risky in a diff, and all of them can quietly change answers. Without a measurement on each change, improvement is an opinion and regression is a surprise.

Build a golden set

A golden set is a versioned collection of test cases kept next to the code. Each case holds a question, the sources that should support the answer and, where it matters, the expected behaviour: an answer, a refusal or an escalation.

  • Start from real questions, taken from users or from the people who know the domain, not invented ones.
  • Include refusals. A system that must decline when it has no source needs cases where declining is the correct result.
  • Add hard and hostile cases: ambiguous questions, near-duplicate documents, attempts to extract personal data.
  • Keep it reviewable. A few dozen well-chosen cases that people have read beat thousands that nobody has.
  • Give it an owner. The set changes through review, like code.

Choose metrics that match the failure you fear

Measure each layer separately, so that a failure points to its cause.

  • Retrieval: did the expected source appear among the retrieved passages? DeepEval offers contextual precision and contextual recall for this.
  • Grounding: is the answer supported by the retrieved passages and do its citations match them? Faithfulness metrics cover the first half, and a deterministic citation check covers the second.
  • Relevance: does the answer address the question? Answer relevancy is the usual metric.
  • Behaviour: does the system decline when it should, and mask personal data when it must? Use exact checks, not a judge.
  • Budgets: latency and cost per request, tracked as thresholds too.

Gate releases

DeepEval lets you write evaluations as tests, so they run in the same pipeline as the rest of your suite. Run the golden set whenever something that affects behaviour changes: a prompt, a model, retrieval settings, a tool manifest.

  • Set thresholds per metric, and start by recording the current scores so that the first gate is “no worse than today”.
  • Expect noise. Model output varies, so compare against a tolerance, or repeat a case and use the average, instead of demanding identical scores.
  • Split the suite. A small, fast set runs on every pull request, and the full set runs before release.
  • Fail loudly. A failed gate should name the cases and metrics that dropped.

Use judges carefully

Many metrics use a model as the judge. That is practical, and it has limits: a judge can drift when its own model changes, and it can be lenient with fluent but wrong answers. Pin the judge model and version, calibrate it against a small set of human-labelled cases, and keep deterministic checks for anything that can be checked exactly.

Close the loop with production traces

Offline evaluation only covers the cases you thought of. Langfuse records traces of real calls, manages prompt versions and holds datasets, which is what you need to learn from production. When a trace shows a bad answer, reproduce it, fix the cause and add the case to the golden set, so that the same failure cannot return unnoticed. Tie each trace to the prompt version that produced it, so a regression can be traced to a change.

In practice

On the ENSO platform I specified evaluation gates in CI built on DeepEval, golden sets and Langfuse tracing, as one layer of a platform that also includes a gateway, retrieval and governed agent actions. The wider architecture is described in the enterprise AI platform reference guide and in the ENSO case study.

Trade-offs and limits

Evaluation has a cost: running a judge model on every case takes time and money, so the suites need to be sized. Golden sets go stale as products and documents change, and a team can overfit to its own set, so refresh it from production. A passing gate shows that known cases still work, not that the system is correct in general, and that is why human review of real conversations stays in the loop.