Join 5,000+ engineers & PMs in mastering AI Evals, Oct 10 cohort 25% off →

  • Blog
  • Notes
  • OSS
  • Teaching
  1. Notes
  2. LLMs
  3. AI evals resources
  4. Evals Flashcards
  • Notes
    • Python Concurrency
    • CUDA Version Management
    • GitHub Actions
    • dbt
    • Docker
    • How to learn
    • Linux
    • pandoc filters
    • programming languages
    • Video Editing
    • Coding Agents
      • Amp
    • fastai
      • Fundamentals
      • Image Classification
      • Data
    • FastHTML
      • Building Annotation Apps with FastHTML
      • Concurrency For Starlette Apps (e.g FastAPI / FastHTML)
    • Jupyter
      • Fix Jupyter CUDA cache
    • K8s
      • Basics
      • Secrets
      • Storage
        • Storage Basics
        • Dynamic Provisioning
      • Scaling
        • ReplicaSets
        • Scaling
      • StatefulSet
      • Jobs & CronJobs
      • Rollouts
      • Helm
        • Creating Helm Charts
        • Testing With Helm
      • Multi-Container Pods
        • Multi-Container Pods
        • Ambassador Sidecars
        • Restart Conditions
        • Sharing Processes in MC Pods
      • Developer tips
      • Pod restart vs. replacement
      • Probes
      • Resource Limits
      • Requesting resources
      • Logging
      • Monitoring
      • Ingress
      • Cluster Components
      • Security
        • Network Security
        • Securing Containers
        • Webhooks
        • Updating a K8s Cluster
        • RBAC
      • Workload Placement
      • Auto Scaling
      • Preemption
      • Random TILs
    • LLMs
      • AI Product Engineering Notes
        • Automating Error Analysis
        • Case Study: Putting Evals Into Production
        • Evals for Data Agents
        • Cut Classification Costs With a Model Cascade
        • An Intro to Multivector Retrieval
        • How to Improve Search Agents
        • How to Optimize Retrieval Embeddings
        • How to Choose an OCR Model
        • Steer AI Writing With Footnotes for Agents
        • Debugging Inference Latency
        • Open Weight Model Economics
        • The Case for Agent Sandboxes
        • Turn Eval Results Into a Better Model
      • Inference
        • Optimizing latency
        • Max Inference Engine
        • vLLM & large models
      • OpenAI
        • Function prompts
      • AI evals resources
        • AI evals: Start here
          • Evals: Getting started
          • Evals: What to test
          • Evals: Can I trust my scores?
          • Evals: Hard-to-evaluate products
          • Evals: Reducing time and cost
        • Inspect AI, An OSS Python Library For LLM Evals
        • Evals Flashcards
        • Evals Memes
      • Fine-tuning
        • Dataset Basics
        • LangChain DocumentLoaders
        • Estimating vRAM
        • Curating LLM data
        • Tokenization Gotchas
        • Template-free axolotl
      • Function Calling
        • Llama-3 Func Calling
      • RAG
        • Stop Saying RAG Is Dead
        • P1: I don’t use RAG, I just retrieve documents
        • P2: Modern IR Evals For RAG
        • P3: Optimizing Retrieval with Reasoning Models
        • P4: Late Interaction Models For RAG
        • P5: RAG with Multiple Representations
        • P6: Context Rot
        • P7: You Don’t Need a Graph DB (Probably)
      • Open Office Hours
        • Evals: Doing Error Analysis Before Writing Tests
        • Multi-Turn Chat Evals
        • Observability in LLM Applications
        • Tame Complexity By Scoping LLM Evals
      • Data Processing
        • How to Process Documents at Scale with LLMs
    • Prompt engineering
      • Course
        • Guidelines for Prompting
        • Iterative Prompt Development
        • Summarizing
        • Inferring
        • Transforming
        • Expanding
        • The Chat Format
    • Quarto
      • Syntax Highlighting
      • Listings from data
      • Merge listings
    • ML Serving
      • TF Serving
        • Basics
        • GPUs & Batching
      • TorchServe
        • Basics
        • Serving Your Own Model
      • FastAPI
    • Web Scraping
      • Browser requests to code
      • Transcribe & Diarize Videos

On this page

  • Background
  • Flashcards


  1. Notes
  2. LLMs
  3. AI evals resources
  4. Evals Flashcards

Evals Flashcards

These downloadable flashcards review the important ideas from our AI Evals course.
Published

November 30, 2025

Background

I created these flashcards to help students learn about evals in our AI Evals course.

If you are new to evals, I recommend starting with the Evals FAQ. For something lighter, check out the Evals Memes.

Flashcards

Click on any image to open it full-size, or use the download button in the corner to save it directly.

1

2

3

4

5

6

7

8

9

10

11

12

13


👉 Want to learn more about AI Evals? Check out our AI Evals course. It’s a live cohort with hands on exercises and office hours. Here is a 25% discount code for readers. 👈