AI EvalsJobs
Hiring?
All roles / Snorkel AI / Research Scientist - Frontier Benchmarks

Research Scientist - Frontier Benchmarks

Evals & BenchmarksRemote4mo ago
Source verified: this exact posting was present on Snorkel AI’s careers feed on . View employer source →

What the posting asks for

Doctorate
Named as preferred, or equivalent experience

Read out of the employer's own description. Absence means the posting does not say, not that the answer is no.

The role

About Snorkel

Snorkel AI is the frontier AI data lab, helping teams build the data and environments behind high-performing frontier and agentic AI. We combine technology with research-driven AI data development to create datasets, benchmarks, evals, and custom solutions for real-world AI systems. Founded out of the Stanford AI Lab in 2019, Snorkel works with leading AI labs and enterprises to move from better data to better outcomes.

Excited to help us redefine how AI is built? Apply to be the newest Snorkeler!

ABOUT THE ROLE

We're looking for a Research Scientist to lead the design of next-generation benchmarks and datasets that push the boundaries of frontier model evaluation. You'll define what "good" looks like across a range of hard tasks, drawing on conversations with customers and academic partners to ground your datasets in real performance gaps. You'll build both the benchmarks that measure those gaps and the training data that helps close them, then partner with delivery, product, and go-to-market to bring what you build into production.

This role is ideal for someone who wants to shape how the field measures progress in frontier AI, and who's energized by doing that work inside a fast-moving, cross-functional startup.

MAIN RESPONSIBILITIES

  • Design state of the art datasets that drive frontier model training and evaluation based on current model performance and academic partnerships
  • Translate benchmark insights into clear, compelling narratives that articulate the ROI of expert-curated data for customer-facing presentations, technical reports, and go-to-market materials.
  • Work cross-functionally with data operations, product, engineering, and strategy to surface research findings that inform the company roadmap.
  • Stay at the frontier of LLM evaluation research and bring best practices into Snorkel's workflows
  • Represent Snorkel's research externally through publications, blog posts, conference talks, and customer engagements that advance the conversation around data-centric AI

PREFERRED QUALIFICATIONS

  • Strong research background in AI/ML evaluation, NLP, or related fields, with a track record of rigorous experimental design — especially around measuring the impact of training and evaluation data on model behavior.
  • Exceptional communication skills — able to present complex technical findings clearly to both technical and non-technical audiences
  • Comfort operating in a fast-moving, cross-functional environment with ambiguous problem spaces
  • Genuine interest in GTM strategy, startup dynamics, and the commercial side of AI data services.
  • Ph.D. in machine learning, NLP, or a related field preferred; equivalent industry or research lab experience considered.

Actual compensation will be determined based on factors including skills, qualifications, experience, and geographic location.

Salary range(s) for this role$200,000—$375,000 USD

Be Your Best at Snorkel

Joining Snorkel AI means becoming part of a company that has market proven solutions, robust funding, and is scaling rapidly—offering a unique combination of stability and the excitement of high growth. As a member of our team, you’ll have meaningful opportunities to shape priorities and initiatives, influence key strategic decisions, and directly impact our ongoing success. Whether you’re looking to deepen your technical expertise, explore leadership opportunities, or learn new skills across multiple functions, you’re fully supported in building your career in an environment designed for growth, learning, and shared success.

Snorkel AI is proud to be an Equal Employment Opportunity employer and is committed to building a team that represents a variety of backgrounds, perspectives, and skills. Snorkel AI embraces diversity and provides equal employment opportunities to all employees and applicants for employment. Snorkel AI prohibits discrimination and harassment of any type on the basis of race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state, or local law. All employment is decided on the basis of qualifications, performance, merit, and business need.

We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.

Published by Snorkel AI on their own careers page and reproduced here unedited. Read it at Snorkel AI.

Apply at Snorkel AI → Applications go directly to Snorkel AI. This board does not sit in between, take a fee from you, or see your application.

What this listing does not tell you

Listed 133 days, which is longer than most. Of the 114 evals and benchmarks roles this board has watched from listing to removal, 81% were gone from their employer's careers page by day 133, and the median came down after 56 days. That is a description of other listings that have already ended, not a prediction about this one: this board records when a listing disappears, never why, and a posting still up is not on a clock it can see.

Snorkel AI has 18 roles open on this board, 12 of them in evals and benchmarks.

Get the weekly AI Evals Jobs briefNew roles and board updates, with published pay where available. This is the general weekly brief. Or browse them all now.

More roles like this

Matched by discipline, title, listed location and work arrangement.

Same discipline: Evals & Benchmarks · Shared listed location

Same discipline: Evals & Benchmarks · Shared listed location · Both list remote work; check location eligibility

Same discipline: Evals & Benchmarks · Shared listed location · Both list remote work; check location eligibility

Same discipline: Evals & Benchmarks · Shared listed location · Both list remote work; check location eligibility

Same discipline: Evals & Benchmarks · Shared listed location · Both list remote work; check location eligibility

Privacy · Terms