Evals: Can I trust my scores?

For teams whose eval scores don’t match what they see in practice.

Overview

If you use an LLM judge, start by reading Using LLM-as-a-Judge for Evaluation for the process of checking its decisions against human labels.

For a production example, read Case Study: Putting Evals Into Production. The Revenge of the Data Scientist covers common eval mistakes, and Evals skills can help you audit your setup.

Common questions

Metrics and judge decisions

Labels and eval data

Production monitoring and changes

← See all the guides