Learn AI evals by solving cases.

The hands-on course for product managers. Review real AI failures, build your eval instincts, and ship AI features you can trust.

Start your first case

Free to start · No credit card · First case in 5 minutes

eval / Helpdesk assistant
0/4 checked

The rule

Returns within 30 days, for unopened items only. Don't promise exceptions.

Trace 01 / 04your call

Customer asks

Can I return this after 45 days?

AI replies

“Of course! We accept returns any time, no questions asked.”

Does this reply pass the rule?

02I opened the box. Refund?queued
03How long do refunds take?queued
04Do you refund shipping?queued
Pass rate—

A sample run. Your call first, then the rest of the suite.

How it works

1. Open a case

Each case is a real-feeling AI product failure, with the business context a PM would actually get.

2. Investigate

Label outputs, spot patterns, pick metrics, and fix the judge that grades your model.

3. Rank up

Earn XP, pass rank exams on fresh scenarios, and get a shareable certificate.

What you'll learn

Error analysis

Read traces, code failures in plain language, and build a taxonomy.

Golden datasets

Coverage, edge cases, and when synthetic data helps or hurts.

Metrics and rubrics

Precision, recall, and rubrics that two humans actually agree on.

LLM-as-judge

Write a judge prompt, then validate it against human labels.

RAG and agent evals

Retrieval quality, faithfulness, tool calls, and trajectories.

Online monitoring

A/B tests, feedback signals, drift, and release gates.

The game layer

  • XP

    Every step you solve moves the needle.

  • Streaks

    One step a day keeps the streak alive.

  • Ranks

    Four ranks, each ending in an exam.

  • Titles

    Unlock titles that show next to your name.

  • Certificates

    Verifiable, shareable, printable.

Your first case is waiting.

A support bot just invented a refund policy. Find out how you'd catch it.

Start your first case