San Francisco, CA · New York City, NY$300k – $400k
Route 01 · I build software
Build the systems that measure models
Evaluation engineering is the work of turning a question about model behaviour into something that runs, repeats and can be trusted. Harnesses, graders, datasets, the infrastructure underneath them, and the analysis that says whether a change actually moved anything.
What these postings actually ask for
Counted from 241 of the 307 open evals & benchmarks roles on this board whose description this board can read, on 2026-09-06. Every figure below is a count of employers' own words. Where a requirement is not stated, that is reported as not stated rather than assumed.
- Named tools. python in 140, pytorch in 69, sql in 30, databricks in 27, jax in 26, kubernetes in 22, tensorflow in 19, rust in 15. Counted where the description names them, which understates any tool a team uses without writing it down.
- Doctorate. 172 of 241 (71%) never raise one. 69 of 241 (29%) mention one, and how many of those require it is not reported here: telling a requirement from an alternative or a description of the team is a reading this board does not yet do reliably, and the page on it explains why the count was withdrawn.
- Years of experience. 38 postings state a number; the median is 4, running from 1 to 7. The other 203 state none.
- On call. 5 of 241 (2%) mention an on-call rotation.
- Security clearance. 3 of 241 (1%) mention one.
- Exposure to distressing material. 4 of 241 (2%) say the work involves it. Worth knowing before an interview rather than after.
- Published pay. 148 of 307 open evals & benchmarks roles advertise a comparable annual USD base band; the rest publish nothing or publish something that is not base pay.
- Remote. 72 of 307 say remote in the listing. Geography restrictions usually sit in the posting itself.
- Engagement. 1 of 307 are contributor work rather than staff requisitions, and are labelled as such on the board.
Suggested by this board, not by an employer
Nothing in this section was asked for by anyone hiring. It is one editor's suggestion for demonstrating the work above, and no posting on this board requires it.
A project
Build an evaluator for a task you understand, and document where it is wrong.
Take a narrow task you can judge yourself. Write the grading criteria down before you look at any output, run two configurations against them, and record every case where your evaluator disagrees with your own judgement. Use public or synthetic examples. The write-up is the deliverable, not the score.
Deliverable: Share the project, your decision criteria, and a short record of what you changed after feedback.
Self-check: Can someone else reproduce your result? Are assumptions and failure cases clear? What evidence would change your conclusion?
A question to ask them
How does this team decide whether a model change is safe to ship?
The answer tells you how the team actually works, which a job description rarely does.
Open now in this track
San Francisco, CA$305k – $385k
San Francisco · New York City$200k – $400k+ equity
San Francisco · New York City$200k – $400k+ equity
Paris · Amsterdam +4Remote listed
Paris · Palo Alto +4Remote listed
US, WA, Seattle
US, WA, Seattle