AI EvalsJobs
Hiring?
All roles / Surge AI / Lead Adversarial Engineer

Lead Adversarial Engineer

Red Team & SafeguardsRemote3mo ago
Source verified: this exact posting was present on Surge AI’s careers feed on . View employer source →

The role

About Us

Our mission is to raise AGI with the richness of human intelligence — curious, witty, imaginative, and full of unexpected brilliance.

Surge was founded by engineers and researchers who dreamed of building the next generation AI. We're building a platform that powers the most powerful models in the world in partnership with companies like Anthropic, Google, Microsoft, and Meta.

At Surge, we believe the path to AGI isn't just about scaling compute—it's about embracing the unlimited ceiling of human intelligence and creativity in the data that shapes these systems. Our platform combines elite human expertise with cutting-edge tools for scalable oversight, from building rich RL environments to conducting rigorous evaluations that go beyond benchmarks. We've run a profitable business from day one without raising venture funding.

The Role

As a Lead Adversarial Engineer, you’ll run end-to-end red-teaming workstreams against frontier models — scoping threat models, designing campaigns, coordinating operators, and synthesizing results into clear risk pictures and decision-ready reports. You’ll orchestrate structured adversarial exercises across modalities and tools, ensuring coverage, reproducibility, and crisp learning loops.

You won’t just find failures — you’ll build the operational engine that repeatedly surfaces them under realistic constraints. This is a role for someone who thrives on program design, loves turning messy attack spaces into disciplined test plans, and can drive cross-functional execution from kickoff to readout.

What You'll Do

  • Stand up a recurring red-team cadence: scoping targets, recruiting operators, defining success criteria, and executing multi-week campaigns

  • Create scenario banks and attack taxonomies; ensuring breadth/depth coverage and tracking families of exploits across versions and contexts

  • Produce executive readouts and issue trackers that distill severity, exploitability, and user harm, with crisp reproduction steps and artifacts

  • Partner with research, product, and ops teams to validate fixes and rerun focused regressions; maintaining dashboards for trendlines and residual risk

What We’re Looking for

  • Red-Team Program Leadership – Experience planning and running adversarial campaigns (jailbreaks, prompt injection, tool abuse), including playbooks, ops cadence, and after-action reviews

  • Methodical Experimentation – Strength in designing scenarios, controls, and metrics; comfort triaging findings and prioritizing next passes based on evidence

  • Stakeholder Command – Ability to brief partners, align on objectives, and translate results into actionable remediation tracks with clear owners and timelines

Published by Surge AI on their own careers page and reproduced here unedited. Read it at Surge AI.

Apply at Surge AI → Applications go directly to Surge AI. This board does not sit in between, take a fee from you, or see your application.

What this listing does not tell you

Listed 88 days, which is longer than most. Of the 23 red team and safeguards roles this board has watched from listing to removal, 70% were gone from their employer's careers page by day 88, and the median came down after 58 days. That is a description of other listings that have already ended, not a prediction about this one: this board records when a listing disappears, never why, and a posting still up is not on a clock it can see.

Pay is not confirmed for this role. Check the employer’s posting for current compensation. Explore published bands from other roles →

Surge AI has 8 roles open on this board.

Get the weekly AI Evals Jobs briefNew roles and board updates, with published pay where available. This is the general weekly brief. Or browse them all now.

More roles like this

Matched by discipline, title, listed location and work arrangement.

Same discipline: Red Team & Safeguards · Both have staff / lead titles · Both list remote work; check location eligibility

Same discipline: Red Team & Safeguards · Both have staff / lead titles · Both list remote work; check location eligibility

Same discipline: Red Team & Safeguards · Both have staff / lead titles

Same discipline: Red Team & Safeguards · Both have staff / lead titles

Same discipline: Red Team & Safeguards · Both have staff / lead titles

Privacy · Terms