AI EvalsJobs
Hiring?
All roles / Mistral AI / Research Engineer - Eval Platform

Research Engineer - Eval Platform

Evals & Benchmarks1d ago
Source verified: this exact posting was present on Mistral AI’s careers feed on . View employer source →

What the posting asks for

Names
Kubernetes, Python, vLLM
Doctorate
Mentioned, without saying whether it is needed

Read out of the employer's own description. Absence means the posting does not say, not that the answer is no.

The role

About Mistral

Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector, co-creating customized AI systems that they can run on their terms.

We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.

The Role

Evaluation is how we decide which models, checkpoints and recipes ship. As a Research Engineer on the Eval Platform team, you will build the infrastructure every science team relies on to measure model quality, and make it reliable, reproducible and fast.

You don't need to have designed benchmarks before. You do need to care about what a score means, and about when a difference between two runs is real.

What you will do

  • Build systems that keep eval results reproducible and comparable over time, as models, benchmarks and code evolve.

  • Run evaluations at scale across our GPU clusters, from model serving to scoring.

  • Make eval results easy to access, explore and trust, through APIs and dashboards that researchers use every day.

  • Catch broken or noisy evals before they mislead research decisions.

  • Support evaluation of agentic, multi-turn and tool-using models.

  • Work closely with researchers to turn new evaluation needs into robust, shared tooling.

What we're looking for

  • Master's or PhD in Computer Science, or equivalent experience.

  • 4+ years building production-grade software, ideally large-scale ML codebases or distributed systems.

  • Excellent Python and strong software-design instincts: testing, code review, CI/CD.

  • Experience running workloads on GPU clusters (Slurm, Kubernetes, Ray or similar).

  • Familiarity with LLM inference and evaluation.

  • A product mindset: researchers are your users.

  • Self-starter, low-ego, collaborative.

Nice to have

  • Experience building or maintaining evaluation harnesses or benchmarks.

  • Hands-on experience with inference engines such as vLLM or SGLang.

  • Experience with agentic or RL environments.

  • Statistics for experimentation: variance estimation, significance testing.

  • Open-source contributions to ML tooling.

What We Offer

We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.

For the most up-to-date details on benefits available in your location, please refer to our Benefits page.

Privacy Policy

Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.

Published by Mistral AI on their own careers page and reproduced here unedited. Read it at Mistral AI.

Apply at Mistral AI → Applications go directly to Mistral AI. This board does not sit in between, take a fee from you, or see your application.

What this listing does not tell you

Listed 1 day, which is recent for this board. Of the 114 evals and benchmarks roles this board has watched from listing to removal, 0% were gone from their employer's careers page by day 1, and the median came down after 56 days. That is a description of other listings that have already ended, not a prediction about this one: this board records when a listing disappears, never why, and a posting still up is not on a clock it can see.

Pay is not confirmed for this role. Check the employer’s posting for current compensation. Explore published bands from other roles →

Mistral AI has 17 roles open on this board, 11 of them in evals and benchmarks.

Get the weekly AI Evals Jobs briefNew roles and board updates, with published pay where available. This is the general weekly brief. Or browse them all now.

More roles like this

Matched by discipline, title, listed location and work arrangement.

Same discipline: Evals & Benchmarks · Shared listed location

Same discipline: Evals & Benchmarks · Shared listed location

Same discipline: Evals & Benchmarks · Shared listed location

Same discipline: Evals & Benchmarks · Shared listed location

Same discipline: Evals & Benchmarks · Shared listed location

Privacy · Terms