Q: Do I need a reference answer or rubric before annotating data?
No. Writing a rubric before you review examples can get in the way.
Let’s get some definitions out of the way:
- A reference answer is an example of a correct response.
- A rubric is a set of criteria for judging a response, such as whether it follows the refund policy.
Both can help, but treat your initial expectations as a starting point that you will revise.
It’s often better to wait until you’ve reviewed some examples before developing a detailed rubric. Reviewers can become so focused on checking each item that they overlook problems outside the rubric. It’s important to give reviewers room to notice things you didn’t anticipate. This change in what you consider good is called “criteria drift”.
For example, let’s say you have a support agent that handles refunds and it escalates refunds to a human per your policy. You might only realize that the process is frustrating for the user after reading a few interactions. Don’t underestimate the degree of criteria drift that will happen as you review examples!
We recommend using error analysis to systematically review examples and decide what might belong in the rubric. This involves writing open-ended notes about what looks wrong, then group similar notes to see which problems recur. See this live demo for a walkthrough.
After doing error analysis, you can write a better rubric informed by user and application behavior. You should periodically do error analysis to make sure your rubric is current.
This article is part of our AI Evals FAQ, a collection of common questions (and answers) about LLM evaluation. View all FAQs or return to the homepage.