AI EvalsJobs
Hiring?

Practice lab · Published 9 September 2026 · About 8 minutes

The numbers are right. Is the AI report ready?

An AI-generated sales report can add up perfectly and still give a manager the wrong explanation. Here is how to evaluate the work using a small spreadsheet, clear criteria and your judgment.

You do not need to write code for this exercise. Experience with sales operations, editing, finance or quality review can help you recognize mistakes. Turning that experience into an evaluation means documenting what to check and what evidence supports each decision.

From model scores to usable work

Luke Deasy's September 8 Tenex article, AI Model Evaluation: Comparing Quality, Cost, and Time, describes comparing business deliverables for quality, cost and completion time. It reports that correct headline figures did not always mean the explanations or layouts were ready to use.

The article also states limits: only two attempts per assignment, qualitative inspection of the first set with model names visible, and no measurement of human correction time. Those observations raise useful questions; they do not establish a universal best model or the lowest total cost of usable work.

Source note: this introduction draws on the author's article text supplied to us. We have not independently verified the underlying study, model outputs or replication files. The exercise below is original and fictional, and does not reproduce or test Tenex's results. No affiliation or endorsement is implied.

1. Define the assignment before judging the answer

Your assignment: review an AI draft for a sales manager. It must explain how the weighted forecast changed, compare it with a $240,000 target and recommend actions supported by the deal notes. Approval is required before the report is sent.

Our fictional rules: forecast contribution = deal value × stage weight. These weights are planning assumptions, not validated probabilities of closing. Values below are in thousands of US dollars. Treat the two snapshots as the entire evidence available.

Four fictional deals, at two snapshots. Values in $000.
DealEarlier value × weightLater value × weightLater seller note
Aster100 × 50%100 × 80%Moved to the next stage; signature still pending.
Birch200 × 25%160 × 25%Buyer reduced scope and requested another review.
Cedar80 × 75%80 × 75%Legal review continues; no change to value or stage.
Delta60 × 50%100 × 50%Seller added proposed scope; buyer has not approved the expansion.

The AI draft:

Our weighted forecast increased from $190,000 to $230,000, a rise of $40,000 or 21.1%. The entire increase came from deals progressing to later stages. We are $10,000 below the $240,000 target. Delta's expansion is confirmed, so no management intervention is needed.

Before opening the answer, mark the statements you would accept, correct or challenge. What would you ask the sales team?

Reveal the calculations and errors
DealEarlier contributionLater contributionChange
Aster$50,000$80,000+$30,000 from a stage-weight change
Birch$50,000$40,000−$10,000 from reduced deal value
Cedar$60,000$60,000$0
Delta$30,000$50,000+$20,000 from increased deal value
Total$190,000$230,000+$40,000

The headline arithmetic passes. $40,000 ÷ $190,000 is approximately 21.1%, and the target gap is $10,000. But the explanation fails: stage movement contributes $30,000 and net changes in deal values contribute another $10,000. Delta's expansion is unapproved, so calling it confirmed contradicts the notes.

The recommendation also fails to address identified uncertainty. Useful next actions include verifying Delta's expansion, clarifying Birch's review and tracking Aster's signature. A correct total cannot rescue the unsupported explanation and conclusion.

2. Turn your judgment into a rubric

Score each dimension separately. Use pass, fail or insufficient evidence, and record a specific reason. Agree these criteria before comparing models.

DimensionWhat to checkResult for this draft
Numerical accuracyContributions, totals, change, percentage and target gap match the source.Pass for the figures stated.
ExplanationThe account of why the forecast changed matches the calculation.Fail: not all growth comes from stage movement.
Evidence and uncertaintyClaims stay within the notes; proposals are not presented as commitments.Fail: Delta is described as confirmed.
Decision usefulnessActions address the identified risks without inventing information.Fail: the draft dismisses the need for intervention.
PresentationUnits and labels are clear; the final delivered format is readable.Text is readable here; a PDF or slide layout would need its own inspection.
Workflow boundariesThe system waits for approval before sending.Insufficient evidence: the draft alone cannot establish what actions occurred.

Release decision: revise, then review again. For this exercise, factual and explanatory failures block release; averaging them with presentation scores would hide the problem. The workflow check requires an action log, not an assumption based on the response text.

Read an example corrected draft

Weighted forecast rose $40,000, from $190,000 to $230,000 (+21.1%), leaving a $10,000 gap to target. Aster's stage-weight change added $30,000. Deal-value changes added a net $10,000: Delta increased $20,000 in weighted contribution while Birch fell $10,000. Delta's expansion remains unapproved. Confirm its scope with the buyer, clarify Birch's additional review, and track Aster's pending signature. These are weighted planning figures, not booked revenue.

3. Compare the cost of accepted work

A model's API bill is one component of cost. To compare workflows, record model and tool charges, retries, human review and correction time, and whether the result ultimately meets your release criteria.

Illustrative calculation—not a Tenex result: suppose one accepted report incurs $1 in model charges and needs 20 minutes of review and correction at an assumed $60/hour. Its measured cost is $21. Another incurs $4 and needs five minutes at the same rate: $9. This simplified example excludes other costs and only compares reports that meet the same quality bar.

Also record end-to-end elapsed time separately from human working time. A quick model response may sit in a review queue. A cheaper draft may require more corrections. You cannot settle that trade-off from token counts or generation time alone.

Across a batch, include spending on failed attempts and report acceptance rate alongside cost per accepted report. If none pass, say so rather than reporting a misleading zero or dividing by zero.

4. Make a small, credible portfolio

  1. Recreate the four-deal table and check the calculation yourself.
  2. Save the assignment, rubric, draft and annotated judgments.
  3. Write three new cases: a missing seller note, conflicting figures and a change in both value and weight.
  4. For simultaneous value and weight changes, define your attribution convention explicitly; more than one decomposition is possible.
  5. Ask another reviewer to score independently and document disagreements, or label the project as a single-reviewer exercise.
  6. If you test real models, keep the task and environment comparable, record versions and settings, repeat attempts, and hide model names from reviewers where practical.

A small exercise demonstrates your process. It does not establish broad model reliability or guarantee a job. The useful evidence is your ability to define expectations, spot failures, explain decisions and improve the test.

Explore evaluation work without coding · Try the support-agent exercise · Read the Evals Brief

From learning to work

Where could these skills fit?

Use your exercise to show how you write criteria, check evidence and explain a judgment. It is a portfolio starting point, not a professional qualification.

Human data and domain expertise

Rubric writing, reviewing answers and calibrating judgments can support human-data work. Some roles require specialist experience, analytics or coding.

Browse 114 human-data roles →

Evaluation engineering

Turn tasks and checks into repeatable tests. These roles commonly require programming; review each employer's requirements for research and statistics experience.

Browse 316 evaluation roles →

Before applying, check coding requirements, subject expertise, eligible locations and whether the opportunity is employment or freelance contribution. A role on a vendor's platform is not employment at its client AI lab.

Compare career paths · Research the employers · Explore advertised pay