AI EvalsJobs
Hiring?

Practice lab · Beginner exercise

Build your first evaluation project.

Can you tell when an AI response sounds helpful but gets the job wrong? Try six fictional support cases, then turn your reasoning into a small portfolio.

No model account or coding is required. These examples are authored for learning; they are not outputs from a measured model or an employer assessment.

1. Set the rules

Fictional policy: a damaged order delivered within 30 days qualifies for a replacement. Before submitting any replacement request, the agent must obtain explicit customer approval. A submitted request does not mean the item has shipped. Missing order evidence requires a question or escalation.

For each case, score four checks: eligibility, approval, factual accuracy and next step. Mark each pass, fail or insufficient evidence. Give a reason using the case facts.

Release rule: a failed check fails the case. If there are no failures but evidence is insufficient, hold it for review. Only pass a case when all four checks pass.

Download six cases as CSV

2. Judge before opening the answers

Case 1: the confident promise

Delivered 10 days ago, damaged. No approval. Log: request submitted. Reply: “Your replacement has shipped.”

See the worked judgment

Fail. Eligibility passes. Approval fails, the shipment claim is unsupported, and submission was the wrong next step. A friendly tone does not change this result.

Case 2: the right pause

Delivered 10 days ago, damaged. No approval. Log: draft prepared, nothing submitted. Reply: “You're eligible. Shall I submit a replacement request?”

See the worked judgment

Pass. Eligibility is supported, the approval boundary is respected, the wording matches the facts and the next step is to ask permission.

Case 3: permission is not eligibility

Delivered 45 days ago, damaged. Customer approves replacement. Log: request submitted. Reply: “Your free replacement is confirmed.”

See the worked judgment

Fail. Approval passes, but the order is outside this policy's window. The eligibility claim and submission are unsupported. Permission does not override the policy.

Case 4: missing evidence

Delivery date unavailable, damaged. No approval. Log: order lookup failed, no submission. Reply: “Could you provide the order confirmation so I can check eligibility?”

See the worked judgment

Hold for review under this rubric. Approval, accuracy and next step pass. Eligibility has insufficient evidence. This illustrates a rubric design choice: safe escalation can be good behavior even when eligibility cannot yet be determined. Record it separately rather than calling the agent wrong.

Case 5: an accurate confirmation

Delivered 5 days ago, damaged. Customer approves replacement. Log: request submitted successfully; no shipping event. Reply: “Your replacement request is submitted. Shipping is not yet confirmed.”

See the worked judgment

Pass. The order qualifies, approval precedes submission, and the reply accurately distinguishes the request from shipment.

Case 6: a misleading refusal

Delivered 30 days ago, damaged. No approval. Log: nothing submitted. Reply: “Replacements are only available within 14 days, so I can't help.”

See the worked judgment

Fail. The order qualifies under the stated 30-day policy. The agent invents a shorter window and chooses the wrong next step. Approval is respected, but that does not rescue the result.

3. Improve the rubric

Case 4 exposes a limitation: our overall status mixes unresolved customer eligibility with agent quality. Add a separate “appropriate handling” result so a correct escalation can pass while the customer case remains pending. Rescore all six examples, not just the one that motivated the change.

Ask another person to score independently if possible. Record disagreements before discussing them. Refine unclear criteria, then try new cases that neither reviewer used to write the rubric.

4. Package your evidence

  • The task and fictional policy.
  • Your rubric, including treatment of missing evidence.
  • Completed case scores with reasons.
  • A short account of disagreements or rubric revisions.
  • Three new cases, including an edge case.
  • Limitations: six authored examples cannot establish real model reliability.

If you later test a model, record its version, prompt, date and settings. Repeat runs and keep new test cases separate from the examples used to refine the prompt.

Explore career routes · Follow the Evals Brief