AI evals: Start here

Start with the row that sounds most like your situation.

Where are you? This sounds like me
I’m new to evals I’ve heard the term, but I’m not sure what evals involve or whether I need them.
I don’t know what to test I’m building an AI product, but I haven’t figured out which failures to measure or what good performance looks like.
I don’t trust my eval scores We have evals, but the scores don’t match our judgment of the outputs, or tests pass while users still encounter problems.
My product feels too hard to evaluate Our outputs are subjective, long, or involve many steps. Even a knowledgeable person has trouble deciding whether they’re right.
Evals take too much time or money We’re spending too much effort reviewing outputs, maintaining tests, or running evaluators.