Q: How are evaluations used differently in CI/CD vs. monitoring production?

LLMs
evals
faq
faq-individual
CI evals use small test datasets, while production evals sample live traces.
Authors

Hamel Husain

Shreya Shankar

Published

June 29, 2025

Modified

September 1, 2026

CI evals protect against known regressions before deployment. Online monitoring find failures in production traffic and estimate how often they occur.

Evals in CI

Test datasets for CI are small (in many cases 100+ examples) and purpose-built. Examples cover core features, regression tests for past bugs, and known edge cases. Since CI tests are run frequently, the cost of each test has to be carefully considered (that’s why you carefully curate the dataset). Favor assertions or other deterministic checks over LLM-as-judge evaluators.

Onnline monitoring for production

For evaluating production traffic, you can sample live traces and run evaluators against them asynchronously. Since you usually lack reference outputs on production data, you might rely more on on more expensive reference-free evaluators like LLM-as-judge. Additionally, track confidence intervals for production metrics. If the lower bound crosses your threshold, investigate further.

Connect the two systems

These two systems are complementary: when production monitoring reveals new failure patterns through error analysis and evals, add representative examples to your CI dataset. This mitigates regressions on new issues.

Here is a visual that helps contrast the approaches.

↩︎ Back to main FAQ


This article is part of our AI Evals FAQ, a collection of common questions (and answers) about LLM evaluation. View all FAQs or return to the homepage.