AI EvalsJobs
Hiring?
All roles / Perplexity / Member of Technical Staff (Machine Learning Engineer, Search & Agents)

Member of Technical Staff (Machine Learning Engineer, Search & Agents)

Evals & Benchmarks3d ago
Source verified: this exact posting was present on Perplexity’s careers feed on . View employer source →

The role

Perplexity is seeking an experienced Machine Learning Engineer to advance how AI systems search, reason, and work together to solve complex problems. Our work spans search and retrieval, LLM post-training, multi-agent training, and the harnesses that make these systems effective.

We control the full stack: the models, the agent harnesses, and the search infrastructure underneath. That gives us the freedom to develop new approaches across all three training models to use search more effectively, designing tools and execution environments around learned behavior, and improving retrieval to support how agents actually work.

Responsibilities

  • Push search and agent quality forward through improvements to models, training data, tools, and system design.

  • Develop LLM post-training methods, including reinforcement learning, to improve reasoning, search, tool use, and task completion.

  • Train and evaluate multi-agent systems, exploring how agents divide work, share information, and coordinate effectively.

  • Design and build agent harnesses around the tools, context management, execution environments, and orchestration that support reliable work over many steps.

  • Improve retrieval and ranking models and the search interfaces agents use to find and assess information.

  • Build datasets, reward signals, and evaluations that expose meaningful failures and guide improvements.

  • Own experiments end to end, from a clear hypothesis to scalable training, deployment, and measurable gains in quality, latency, and cost.

  • Collaborate with AI, Search, Infrastructure, Data, and Product teams to bring new capabilities into production.

Qualifications

  • A strong track record of building and shipping ML systems, with deep experience in one or more of LLM post-training, reinforcement learning, search and retrieval, or agent systems.

  • Strong software engineering skills and the ability to work across model training, experimentation infrastructure, and production systems.

  • Experience designing rigorous evaluations, diagnosing failures, and translating experimental results into practical improvements.

  • Comfort with open-ended problems that require both research judgment and hands-on engineering.

  • A strong sense of ownership, curiosity, and the drive to carry an idea through to a working system.

Nice to have

  • Experience with training models to use tools or complete tasks over many steps.

  • Multi-agent training, coordination, or evaluation.

  • Building agent harnesses, distributed training systems, or scalable inference infrastructure.

  • Large-scale retrieval, ranking, or recommendation systems.

Published by Perplexity on their own careers page and reproduced here unedited. Read it at Perplexity.

Apply at Perplexity → Applications go directly to Perplexity. This board does not sit in between, take a fee from you, or see your application.

What this listing does not tell you

Listed 3 days, which is recent for this board. Of the 114 evals and benchmarks roles this board has watched from listing to removal, 3% were gone from their employer's careers page by day 3, and the median came down after 56 days. That is a description of other listings that have already ended, not a prediction about this one: this board records when a listing disappears, never why, and a posting still up is not on a clock it can see.

Pay is not confirmed for this role. Check the employer’s posting for current compensation. Explore published bands from other roles →

Perplexity has 13 roles open on this board, 9 of them in evals and benchmarks. 8 of those do publish a band, which you can compare; pay can differ by role and location.

Get the weekly AI Evals Jobs briefNew roles and board updates, with published pay where available. This is the general weekly brief. Or browse them all now.

More roles like this

Matched by discipline, title, listed location and work arrangement.

Same discipline: Evals & Benchmarks · Both have staff / lead titles · Shared listed location

Same discipline: Evals & Benchmarks · Both have staff / lead titles · Shared listed location

Same discipline: Evals & Benchmarks · Both have staff / lead titles · Shared listed location

Same discipline: Evals & Benchmarks · Both have staff / lead titles · Shared listed location

Same discipline: Evals & Benchmarks · Both have staff / lead titles · Shared listed location

Privacy · Terms