AI EvalsJobs
Hiring?
All roles / Ways in / I run programmes or operations

Route 04 · I run programmes or operations

Make human evaluation work at scale

Evaluation programmes need people who can coordinate expert reviewers, define quality gates, hold a delivery schedule, and turn a research need into something repeatable. In a posting, the signal to look for is ownership of the evaluation process rather than support of it.

What these postings actually ask for

Counted from 91 of the 105 open human data roles on this board whose description this board can read, on 2026-09-06. Every figure below is a count of employers' own words. Where a requirement is not stated, that is reported as not stated rather than assumed.

  • Named tools. python in 23, sql in 15, label studio in 11, spark in 4, databricks in 3, inspect in 3, kubernetes in 2, typescript in 2. Counted where the description names them, which understates any tool a team uses without writing it down.
  • Doctorate. 86 of 91 (95%) never raise one. 5 of 91 (5%) mention one, and how many of those require it is not reported here: telling a requirement from an alternative or a description of the team is a reading this board does not yet do reliably, and the page on it explains why the count was withdrawn.
  • Years of experience. 22 postings state a number; the median is 4, running from 1 to 10. The other 69 state none.
  • Exposure to distressing material. 2 of 91 (2%) say the work involves it. Worth knowing before an interview rather than after.
  • Published pay. 29 of 105 open human data roles advertise a comparable annual USD base band; the rest publish nothing or publish something that is not base pay.
  • Remote. 44 of 105 say remote in the listing. Geography restrictions usually sit in the posting itself.
  • Engagement. 34 of 105 are contributor work rather than staff requisitions, and are labelled as such on the board.

See the 105 open human data roles →

Suggested by this board, not by an employer

Nothing in this section was asked for by anyone hiring. It is one editor's suggestion for demonstrating the work above, and no posting on this board requires it.

A project

Design an evaluation pilot end to end, on one page.

Plan the review of two hundred model responses. Specify how reviewers are selected and calibrated, how you sample, how disagreements are resolved, what the schedule is, and what acceptance looks like. Say what you would measure in the first week and what would make you stop.

Deliverable: Share the project, your decision criteria, and a short record of what you changed after feedback.

Self-check: Can someone else reproduce your result? Are assumptions and failure cases clear? What evidence would change your conclusion?

A question to ask them

What decisions would I own, and how is evaluation quality measured here?

The answer tells you how the team actually works, which a job description rarely does.

Open now in this track

Remote InternationalRemote listed$35–$75/hr

Human Data & AnnotationContributorNew3d ago

All 105 →