AI EvalsJobs
Hiring?
All roles / Mistral AI / AI Scientist - Agentic Engineering

AI Scientist - Agentic Engineering

Evals & Benchmarks5w ago
Source verified: this exact posting was present on Mistral AI’s careers feed on . View employer source →

What the posting asks for

Names
Python
Doctorate
Not mentioned

Read out of the employer's own description. Absence means the posting does not say, not that the answer is no.

The role

About Mistral

Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector, co-creating customized AI systems that they can run on their terms.

We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.

The Role

Mistral is looking for AI Scientists with deep ML expertise and hands-on engineering experience to expand what our agentic tools can do across the engineering lifecycle — CAE (CFD, FEA, etc.) and EDA/Semi.

Working within the AI4Engineering Science team, your core work is building the pre and post-training data for Mistral's LLMs to reason about and execute real engineering tasks. Because Mistral trains its own frontier LLMs, the data and verifiers you design ship directly into models you can hold, a rare position, and the core of the job.

Alongside this, you'll help shape the agent architectures and harness that let these models operate reliably inside multi-step engineering workflows, not just answer isolated questions.

You'll work closely with domain experts across CAE and EDA or other domains to ground this work in how engineers actually work, and with the broader research team to translate that domain grounding into training signal and evaluation benchmarks that measure genuine task competence.

This is early-stage work, and that's the point: you'd be joining at the foundation, shaping the data, verifiers, and agent scaffolding that decide whether these systems become reliable or stay demo-grade. There's no inherited playbook, you'll help define what good looks like, and your work will set the direction the team builds on rather than extend an existing one.

What you will do

  • Design pretraining, SFT, and RL data for engineering tasks across CAD, CAE, and semiconductor/EDA

  • Define verifiers and evaluation criteria that capture what "correct" and "high-quality" actually mean for each engineering task, beyond surface-level plausibility

  • Design and improve agent architectures and harnesses: how models plan, call tools, recover from errors, and chain steps together across long-horizon engineering workflows

  • Build evaluation benchmarks and diagnostic tooling to identify where models fail on engineering tasks, and trace those failures back to gaps in data, reward design, or agent scaffolding

  • Collaborate with domain experts across CAE and EDA or other domains (and the science and solutions teams more broadly) to identify which engineering workflows are highest-value to target next

  • Contribute to Mistral's broader pre and post-training research, sharing findings and methodology across the science organization

What we're looking for

  • Fluent English with excellent communication skills, able to explain technical ML and engineering concepts to both engineering and non-technical audiences

  • Deep, hands-on machine learning expertise, particularly LLM development

  • Demonstrated experience running, debugging, and validating real engineering workflows in at least one of CAD, CAE, semiconductor simulation, or EDA

  • You write clean, readable Python code and are comfortable in Linux/HPC environments

  • Self-directed, you don't need detailed roadmaps to make progress

  • Low-ego, collaborative, and eager to learn at the intersection of engineering and ML

It would be great if you

  • Have experience building or fine-tuning agentic systems (tool use, multi-step planning, agent orchestration frameworks)

  • Have experience with reward modeling, RLHF/RLAIF/RLVR, or preference-based training

  • Have industrial or academic experience with CAE or EDA tools (e.g. SolidWorks, CATIA, Fluent, Abaqus, LS-DYNA, STAR-CCM+, Cadence/Synopsys/Siemens EDA tools)

  • Have contributed to a large open-source or industry codebase

  • Have publications in engineering and ML venues (NeurIPS, ICLR, JFM, AIAA, etc.)

What We Offer

We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.

For the most up-to-date details on benefits available in your location, please refer to our Benefits page.

Privacy Policy

Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.

Published by Mistral AI on their own careers page and reproduced here unedited. Read it at Mistral AI.

Apply at Mistral AI → Applications go directly to Mistral AI. This board does not sit in between, take a fee from you, or see your application.

What this listing does not tell you

Listed 35 days. Of the 114 evals and benchmarks roles this board has watched from listing to removal, 36% were gone from their employer's careers page by day 35, and the median came down after 56 days. That is a description of other listings that have already ended, not a prediction about this one: this board records when a listing disappears, never why, and a posting still up is not on a clock it can see.

Pay is not confirmed for this role. Check the employer’s posting for current compensation. Explore published bands from other roles →

Mistral AI has 17 roles open on this board, 11 of them in evals and benchmarks.

Get the weekly AI Evals Jobs briefNew roles and board updates, with published pay where available. This is the general weekly brief. Or browse them all now.

More roles like this

Matched by discipline, title, listed location and work arrangement.

Same discipline: Evals & Benchmarks · Shared listed location

Same discipline: Evals & Benchmarks · Shared listed location

Same discipline: Evals & Benchmarks · Shared listed location

Same discipline: Evals & Benchmarks · Shared listed location

Same discipline: Evals & Benchmarks · Shared listed location

Privacy · Terms