AI EvalsJobs
Hiring?
All roles / Ways in / I work in ML or research

Route 02 · I work in ML or research

Make model progress measurable

Post-training and evaluation research sits between experimental design and model behaviour. Depending on the team it covers preference data, reward models, reasoning benchmarks, and the analysis that separates a real capability change from a measurement artefact.

What these postings actually ask for

Counted from 135 of the 156 open post-training & rl roles on this board whose description this board can read, on 2026-09-06. Every figure below is a count of employers' own words. Where a requirement is not stated, that is reported as not stated rather than assumed.

  • Named tools. python in 63, pytorch in 41, jax in 20, kubernetes in 20, tensorflow in 16, vllm in 11, databricks in 9, typescript in 9. Counted where the description names them, which understates any tool a team uses without writing it down.
  • Doctorate. 88 of 135 (65%) never raise one. 47 of 135 (35%) mention one, and how many of those require it is not reported here: telling a requirement from an alternative or a description of the team is a reading this board does not yet do reliably, and the page on it explains why the count was withdrawn.
  • Years of experience. 16 postings state a number; the median is 5, running from 2 to 10. The other 119 state none.
  • On call. 2 of 135 (1%) mention an on-call rotation.
  • Published pay. 95 of 156 open post-training & rl roles advertise a comparable annual USD base band; the rest publish nothing or publish something that is not base pay.
  • Remote. 39 of 156 say remote in the listing. Geography restrictions usually sit in the posting itself.

See the 156 open post-training & rl roles →

Suggested by this board, not by an employer

Nothing in this section was asked for by anyone hiring. It is one editor's suggestion for demonstrating the work above, and no posting on this board requires it.

A project

Test whether a benchmark measures what you think it does.

Take a small public benchmark. Vary prompt wording or task format, measure how much the score moves, and inspect the failures rather than the average. Report uncertainty and possible contamination instead of concluding that one model is better.

Deliverable: Share the project, your decision criteria, and a short record of what you changed after feedback.

Self-check: Can someone else reproduce your result? Are assumptions and failure cases clear? What evidence would change your conclusion?

A question to ask them

What evidence would make you change your mind about this result?

The answer tells you how the team actually works, which a job description rarely does.

Open now in this track

All 156 →