Q: What should I do when my “gold” eval dataset becomes stale?

LLMs
evals
faq
faq-individual
Refresh eval datasets with regular error analysis as your product and users change.
Authors

Hamel Husain

Shreya Shankar

Published

September 15, 2026

Modified

September 15, 2026

Eval datasets naturally get stale as your product and users change. Use regular error analysis to find new problems and update your examples or reference answers. How often you review depends on your use case and how quickly your product or usage changes.

Like unit tests, evals can catch problems that return after a fix. However, evals often cost considerably more than unit tests to maintain and run. Therefore, you should weigh each eval’s cost against the value of its signals. If everything keeps passing, this is a sign that the eval is no longer useful and should be retired or run less often.

As your eval set changes, its scores may no longer be directly comparable with older scores. That is ok! One purpose of evals are to provide you with challenges you can hill climb against. These challenges should change as your product evolves to help you keep improving.

For tracking progress with metrics over longer time horizons, it’s often better to use product metrics in addition to evals. Examples of product metrics include: churn, active users, revenue, etc.

↩︎ Back to main FAQ


This article is part of our AI Evals FAQ, a collection of common questions (and answers) about LLM evaluation. View all FAQs or return to the homepage.