Palo Alto, CA$180k – $440k
Member of Technical Staff - Post-Training and RL
The role
SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.
ABOUT THE ROLE:
- You will work on the most critical post-training and reinforcement learning challenges at any given time — including reward modeling, preference optimization (RLHF/DPO), and RL for improving reasoning, truthfulness, and real-world capabilities.
- You will get clarity on your first project before an offer.
BASIC QUALIFICATIONS:
- You believe truth-seeking AI is the most important and challenging problem.
- You are obsessed about building incredibly useful models through post-training and RL techniques.
- You are a power user of AI models and eager to push the boundaries of what’s possible with reinforcement learning and alignment methods.
- If you previously worked on post-training, RLHF, or trained models used by millions of people it’s a big plus, but relevant experience is not required.
- You take pride in your work and thrive in meritocratic environments.
COMPENSATION AND BENEFITS:
$180,000 - $600,000 USD
Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.
SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.
Published by xAI on their own careers page and reproduced here unedited. Read it at xAI.
Apply at xAI → Applications go directly to xAI. This board does not sit in between, take a fee from you, or see your application.
What this listing does not tell you
Listed 163 days, which is longer than most. Of the 40 post-training and RL roles this board has watched from listing to removal, 65% were gone from their employer's careers page by day 163, and the median came down after 110 days. That is a description of other listings that have already ended, not a prediction about this one: this board records when a listing disappears, never why, and a posting still up is not on a clock it can see.
xAI has 39 roles open on this board, 4 of them in post-training and RL.
Free. One email on Thursdays.
More roles like this
Matched by discipline, title, listed location and work arrangement.
Same discipline: Post-Training & RL · Both have staff / lead titles · Shared listed location
Palo Alto, CA$180k – $440k
Same discipline: Post-Training & RL · Both have staff / lead titles · Shared listed location
London$250k – $535k+ equity
Same discipline: Post-Training & RL · Both have staff / lead titles
San Francisco$150k – $300k+ equity
Same discipline: Post-Training & RL · Both have staff / lead titles
San Francisco, CA · London +1
Same discipline: Post-Training & RL · Both have staff / lead titles