AI EvalsJobs
Hiring?
All roles / Surge AI / Research Scientist

Research Scientist

Evals & BenchmarksRemote3mo ago
Source verified: this exact posting was present on Surge AI’s careers feed on . View employer source →

The role

About Us

Our mission is to raise AGI with the richness of human intelligence — curious, witty, imaginative, and full of unexpected brilliance.

Surge was founded by engineers and researchers who dreamed of building the next generation AI. We're building a platform that powers the most powerful models in the world in partnership with companies like Anthropic, Google, Microsoft, and Meta.

At Surge, we believe the path to AGI isn't just about scaling compute—it's about embracing the unlimited ceiling of human intelligence and creativity in the data that shapes these systems. Our platform combines elite human expertise with cutting-edge tools for scalable oversight, from building rich RL environments to conducting rigorous evaluations that go beyond benchmarks. We've run a profitable business from day one without raising venture funding.

The Role

As a Research Scientist focused on data, you’ll work at the cutting edge of LLM development by designing the datasets that shape how frontier models behave. You’ll partner directly with AI research teams to experiment with new data collection strategies, evaluate dataset quality, and uncover insights that improve alignment, safety, and model performance.

This is a research-driven, impact-heavy role: you’ll have the opportunity to test ideas quickly in live production environments, shape data-centric methodologies for some of the world’s top AI labs, and help define how high-quality data fuels the next generation of AI systems.

What We’re Looking for

  • Deep Curiosity About Data – Obsession with understanding how the structure, quality, and selection of data influence LLM performance

  • Bias Toward Insight and Iteration – You think critically but move fast, designing lightweight experiments to generate actionable results

  • Desire to Shape AI Development – Drive to build the foundational data layer that defines how the next generation of AI systems behave

What You'll Do

  • Research data collection strategies and designing high-impact data slices that uncover model failure modes

  • Model annotator behavior and designing experiments to optimize instruction clarity and reward signal reliability

  • Develop metrics and frameworks for evaluating dataset quality, diversity, and impact on downstream model alignment

Published by Surge AI on their own careers page and reproduced here unedited. Read it at Surge AI.

Apply at Surge AI → Applications go directly to Surge AI. This board does not sit in between, take a fee from you, or see your application.

What this listing does not tell you

Listed 88 days, which is longer than most. Of the 114 evals and benchmarks roles this board has watched from listing to removal, 69% were gone from their employer's careers page by day 88, and the median came down after 56 days. That is a description of other listings that have already ended, not a prediction about this one: this board records when a listing disappears, never why, and a posting still up is not on a clock it can see.

Pay is not confirmed for this role. Check the employer’s posting for current compensation. Explore published bands from other roles →

Surge AI has 8 roles open on this board, 4 of them in evals and benchmarks.

Get the weekly AI Evals Jobs briefNew roles and board updates, with published pay where available. This is the general weekly brief. Or browse them all now.

More roles like this

Matched by discipline, title, listed location and work arrangement.

United States - RemoteRemote listed

Evals & BenchmarksFull timeNew1w ago

Same discipline: Evals & Benchmarks · Shared listed location · Both list remote work; check location eligibility

Same discipline: Evals & Benchmarks · Shared listed location · Both list remote work; check location eligibility

Same discipline: Evals & Benchmarks · Both list remote work; check location eligibility

Same discipline: Evals & Benchmarks · Both list remote work; check location eligibility

Same discipline: Evals & Benchmarks · Both list remote work; check location eligibility

Privacy · Terms