AI EvalsJobs
Hiring?
All roles / Surge AI / RL Environments Architect

RL Environments Architect

Post-Training & RLRemote3mo ago
Source verified: this exact posting was present on Surge AI’s careers feed on . View employer source →

The role

About Us

Our mission is to raise AGI with the richness of human intelligence — curious, witty, imaginative, and full of unexpected brilliance.

Surge was founded by engineers and researchers who dreamed of building the next generation AI. We're building a platform that powers the most powerful models in the world in partnership with companies like Anthropic, Google, Microsoft, and Meta.

At Surge, we believe the path to AGI isn't just about scaling compute—it's about embracing the unlimited ceiling of human intelligence and creativity in the data that shapes these systems. Our platform combines elite human expertise with cutting-edge tools for scalable oversight, from building rich RL environments to conducting rigorous evaluations that go beyond benchmarks. We've run a profitable business from day one without raising venture funding.

The Role

As an RL Environments Architect, you’ll design, instrument, and govern the simulated worlds where agents learn — from compact task microcosms to multi-agent, tool-using ecosystems. You’ll define the primitives, reward structures, interfaces, and telemetry that let us stress-test emerging capabilities while keeping training signals faithful, stable, and scalable.

Not only will you build environments, you’ll craft standards for data quality and reproducibility across large-scale agent gyms. This is a role for someone who sweats the details of simulation fidelity, thinks in terms of coverage and failure surfaces, and loves turning messy real-world phenomena into learnable curricula. Your work will form the backbone for safe, rapid progress in agentic systems.

What You'll Do

  • Architect a modular environment framework with clear APIs, curriculum scaffolds, and configurable reward/termination schemas

  • Establish quality bars: coverage metrics, invariance checks, and trace audits for environment outputs and agent experience buffers

  • Instrument rich telemetry for episode rollouts; mitigating reward hacking, mode collapse, and exploitable loopholes

  • Partner with researchers to translate real-world tasks into robust simulations, including synthetic data generators and evaluation suites

What We’re Looking for

  • Simulation & Systems Depth – Experience building RL environments or simulators (e.g., custom physics, multi-agent, tool APIs) with an eye for determinism, performance, and observability

  • Data Quality Leadership – Strong instincts for designing reward functions, scenario taxonomies, and QA pipelines that keep signals aligned and drift-free

  • Builder’s Mindset – Comfort collaborating across research and engineering to ship pragmatic, testable environments that evolve with model capabilities

Published by Surge AI on their own careers page and reproduced here unedited. Read it at Surge AI.

Apply at Surge AI → Applications go directly to Surge AI. This board does not sit in between, take a fee from you, or see your application.

What this listing does not tell you

Listed 88 days. Of the 40 post-training and RL roles this board has watched from listing to removal, 43% were gone from their employer's careers page by day 88, and the median came down after 110 days. That is a description of other listings that have already ended, not a prediction about this one: this board records when a listing disappears, never why, and a posting still up is not on a clock it can see.

Pay is not confirmed for this role. Check the employer’s posting for current compensation. Explore published bands from other roles →

Surge AI has 8 roles open on this board.

Get the weekly AI Evals Jobs briefNew roles and board updates, with published pay where available. This is the general weekly brief. Or browse them all now.

More roles like this

Matched by discipline, title, listed location and work arrangement.

Same discipline: Post-Training & RL · Both list remote work; check location eligibility

Same discipline: Post-Training & RL · Both list remote work; check location eligibility

Same discipline: Post-Training & RL · Both list remote work; check location eligibility

Same discipline: Post-Training & RL

Same discipline: Post-Training & RL · Both list remote work; check location eligibility

Privacy · Terms