AI EvalsJobs
Hiring?
All roles / NVIDIA / Senior Technical Program Manager - LLM Safety

Senior Technical Program Manager - LLM Safety

Evals & Benchmarks5w ago
Source verified: this exact posting was present on NVIDIA’s careers feed on . View employer source →

What the posting asks for

Experience
10+ years asked for

Read out of the employer's own description. Absence means the posting does not say, not that the answer is no.

The role

At NVIDIA, we redefine what’s possible in digital imaging, personal computer gaming, and high-performance computing. Now, we lead the charge into AI’s unlimited potential. Our GPUs act as the brains of computers, robots, and self-driving cars. They enable these machines to understand and interact with the world in new ways. This is your chance to join a legacy of innovation and excellence. Work with the world’s best talent in a diverse and encouraging environment. Join us and make a lasting impact on the world!

This is an ambitious opportunity to work at the forefront of AI safety and contribute to groundbreaking advancements in technology.

What You’ll Be Doing:

  • Lead cross-functional planning and execution for evaluation identification, selection and execution. With a focus on agentic-safety evaluations, including multi-turn tool calling, unsafe tool use, unintended actions, excessive autonomy, and multi-step failures.

  • Track execution, results for all evals, analysis, and the mitigation and research plans for improvement that will stem from results analysis. Translate evaluation findings into prioritized mitigation plans with accountable owners, committed dates, and measurable closure criteria

  • Lead end-to-end safety planning and execution for Nemotron models, including program scope, dependencies, risks, resources, and release-readiness criteria.

  • Translate technical safety priorities into executable programs covering LLM security, frontier risks, agentic safety, hallucinations

  • Establish and track multi-turn, multi-modal, multi-lingual, long-context, and reasoning model evaluations across relevant use cases, domains, languages, modalities, and model releases.

  • Establish governance for reviewing safety findings, assigning severity, determining release impact, and escalating unresolved risks. Ensure required safety evidence, approvals, exceptions, and risk-acceptance decisions are documented before release. Identify program gaps and critical dependencies early and drive decisions, escalations, recovery plans, and corrective actions.

  • Build (with AI for Code/AI for Work) and maintain dashboards and executive reporting for evaluation coverage, critical findings, mitigation progress, residual risk, and release readiness.

What we need to see:

  • Bachelor’s degree or equivalent experience in computer science, engineering, data science, or a related technical field.

  • 10+ years of experience in technical program management, engineering, product development, technical operations, or a similar area.

  • Strong understanding of LLM architecture, frameworks (e.g., OpenAI, Anthropic, Hugging Face), and model evaluation.

  • Familiarity with LLM development, post-training, inference, tool calling, evaluation datasets, and model-release lifecycles.

  • Experience directing programs related to AI/ML development, model evaluation, agentic safety, security, content safety, hallucinations, and/or production releases.

  • Familiarity with AI safety risks such as hallucinations, timely injection attacks, unsafe tool use, data poisoning, model manipulation, and unintended agent behavior.

  • Experience supporting model-safety evaluations, Red Teaming, adversarial testing, security assessments, or responsible-AI programs.

  • Ability to interpret technical evaluation findings and communicate their product, schedule, and release implications to leadership.

Ways to stand out from the crowd:

  • Experience with LLM Agentic Safety, Security, Red Teaming, Hallucinations, adversarial assessment, or Responsible AI.

  • Experience defining evaluation datasets, rubrics, benchmarks, and release gates.

  • Ability to use evaluation results and data to make clear release recommendations and communicate residual risk to leadership.

  • Experience crafting or operationalizing multi-turn, hallucination, adversarial, or agentic-safety evaluations.

With competitive salaries and a generous benefits package, NVIDIA is widely considered to be one of the technology industry's most desirable employers. We have some of the most forward-thinking and versatile people in the world working with us, and our engineering teams are growing fast in some of the most impactful fields of our generation: Autonomous Vehicles or Robotics. If you're a creative Senior Technical Program Manager who embraces autonomy and shares our passion for technology, we want to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 168,000 USD - 258,750 USD for Level 4, and 200,000 USD - 322,000 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 5, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Published by NVIDIA on their own careers page and reproduced here unedited. Read it at NVIDIA.

Apply at NVIDIA → Applications go directly to NVIDIA. This board does not sit in between, take a fee from you, or see your application.

What this listing does not tell you

Listed 38 days. Of the 114 evals and benchmarks roles this board has watched from listing to removal, 39% were gone from their employer's careers page by day 38, and the median came down after 56 days. That is a description of other listings that have already ended, not a prediction about this one: this board records when a listing disappears, never why, and a posting still up is not on a clock it can see.

NVIDIA has 22 roles open on this board, 7 of them in evals and benchmarks.

Get the weekly AI Evals Jobs briefNew roles and board updates, with published pay where available. This is the general weekly brief. Or browse them all now.

More roles like this

Matched by discipline, title, listed location and work arrangement.

Same discipline: Evals & Benchmarks · Both have management titles

Same discipline: Evals & Benchmarks · Both have management titles

Same discipline: Evals & Benchmarks · Both have management titles

Same discipline: Evals & Benchmarks · Both have management titles

Same discipline: Evals & Benchmarks · Both have management titles

Privacy · Terms