AI EvalsJobs
Hiring?
All roles / Who’s hiring

Week of 2026-09-07 · Edition 01

Who’s hiring this week

Five ways into AI evaluation and safety

This week’s picks span frontier-model safety, production safeguards, scientific benchmarks and cybersecurity red teaming. Each vacancy was checked on the employer’s own site on 9 September. “This week” means open when checked—not necessarily first advertised this week.

1. Vijil — Research Scientist / Senior Research Scientist, Fundamental Research

Team: Tim G. J. Rudner’s Fundamental Research team

Study failures in AI agents, multi-agent robustness and scalable oversight, then turn research into tools that improve deployed systems.

Good fit: Researchers with a strong publication record and a PhD or equivalent industry research experience.

Advertised pay: $150,000–$250,000 + equity + benefits as advertised. The page does not explicitly label currency or pay period; confirm the applicable band for your location.

Location: Menlo Park, CA (Remote); New York; Toronto. Confirm remote eligibility for your location.

Official vacancy and application →

Hiring-team signal: Tim G. J. Rudner, Chief Scientist. Earlier hiring announcement, shown as four months old when checked. The current vacancy names the same team and leader; this is not a new announcement this week.

2. OpenAI — Researcher, Agent Safety, Training and Evaluations

Team: Agent Safety / Safety Systems

Turn real agent failures into training data and repeatable evaluations, then test interventions that reduce harmful or misaligned actions.

Good fit: Strong research or ML engineers; prior safety or alignment experience is explicitly not required.

Advertised pay: USD $380,000–$500,000 annual base, plus equity; eligible employees may receive performance bonuses.

Location: San Francisco; hybrid, three office days per week. Relocation assistance offered.

Official vacancy and application → · View and save on AI Evals Jobs

Selected from the official vacancy; no exact hiring-manager post verified.

3. Decagon — Research Engineer, Safety

Team: Research team / Decagon Labs

Build adversarial evaluations and production safeguards against prompt injection, unsafe tool use, data disclosure and unsupported commitments by customer-service agents.

Good fit: Engineers with 2+ years in AI/ML, research or AI safety and hands-on model evaluation, post-training or deployment experience.

Advertised pay: USD $200,000–$400,000 annual base, plus equity. Final compensation varies by location and experience.

Location: San Francisco or New York City; in-office.

Official vacancy and application → · View and save on AI Evals Jobs

Selected from the official vacancy; no exact hiring-manager post verified.

Team background: Team background by Max Lu, 24 March 2026. This is a team-level introduction, not an exact-role hiring-manager post.

4. Edison Scientific — Scientific Evals

Team: Science Research

Design biology benchmarks that test whether AI agents can do useful scientific research; analyze failures and refine tasks, grading and training data.

Good fit: Graduate-level biology or computational-biology training, substantial research experience and working Python skills; wet- and dry-lab experience is especially relevant.

Advertised pay: Advertised annual salary range: USD $160,000–$220,000, plus equity.

Location: San Francisco; full-time, on-site.

Official vacancy and application → · View and save on AI Evals Jobs

Selected from the official vacancy; no exact hiring-manager post verified.

5. Handshake AI — AI Red Teamer, Cybersecurity (Remote)

Team: Handshake AI cybersecurity evaluation

Test whether model-generated security outputs would meaningfully help an attacker, score the risk and document findings that improve model safeguards.

Good fit: Experienced security practitioners with code-analysis skills and hands-on LLM experience.

Advertised pay: USD $65–$125 per hour; contract. Flexible and part-time availability is mentioned, but paid hours are not guaranteed.

Location: Remote within the United States, or Seattle.

Official vacancy and application → · View and save on AI Evals Jobs

Selected from the official vacancy; no exact hiring-manager post verified.

How to read this edition. Listings were open when checked; availability can change. Advertised compensation is not an offer. Hourly contracts are separate from annual salary bands. This is an editorial selection, not a count of newly created jobs or hiring growth.

Browse all roles · Compare published pay

Get the next hiring edition.A curated selection of evaluation and safety roles, with published pay and official application links.