US, WA, Bellevue
Route 02 · I work in ML or research
Make model progress measurable
Post-training and evaluation research sits between experimental design and model behaviour. Depending on the team it covers preference data, reward models, reasoning benchmarks, and the analysis that separates a real capability change from a measurement artefact.
What these postings actually ask for
Counted from 135 of the 156 open post-training & rl roles on this board whose description this board can read, on 2026-09-06. Every figure below is a count of employers' own words. Where a requirement is not stated, that is reported as not stated rather than assumed.
- Named tools. python in 63, pytorch in 41, jax in 20, kubernetes in 20, tensorflow in 16, vllm in 11, databricks in 9, typescript in 9. Counted where the description names them, which understates any tool a team uses without writing it down.
- Doctorate. 88 of 135 (65%) never raise one. 47 of 135 (35%) mention one, and how many of those require it is not reported here: telling a requirement from an alternative or a description of the team is a reading this board does not yet do reliably, and the page on it explains why the count was withdrawn.
- Years of experience. 16 postings state a number; the median is 5, running from 2 to 10. The other 119 state none.
- On call. 2 of 135 (1%) mention an on-call rotation.
- Published pay. 95 of 156 open post-training & rl roles advertise a comparable annual USD base band; the rest publish nothing or publish something that is not base pay.
- Remote. 39 of 156 say remote in the listing. Geography restrictions usually sit in the posting itself.
Suggested by this board, not by an employer
Nothing in this section was asked for by anyone hiring. It is one editor's suggestion for demonstrating the work above, and no posting on this board requires it.
A project
Test whether a benchmark measures what you think it does.
Take a small public benchmark. Vary prompt wording or task format, measure how much the score moves, and inspect the failures rather than the average. Report uncertainty and possible contamination instead of concluding that one model is better.
Deliverable: Share the project, your decision criteria, and a short record of what you changed after feedback.
Self-check: Can someone else reproduce your result? Are assumptions and failure cases clear? What evidence would change your conclusion?
A question to ask them
What evidence would make you change your mind about this result?
The answer tells you how the team actually works, which a job description rarely does.
Open now in this track
San FranciscoRemote listed$350k – $450k
San Francisco, CA · New York, NY$252k – $315k
US, CA, Sunnyvale
US, CA, Santa Clara
US, CA, Sunnyvale
San Francisco · New York$350k – $475k
San Francisco, CA · New York City, NY$500k – $850k