IN, KA, Bengaluru
Manager - Model Evaluation
What the posting asks for
- Names
- Python
- Doctorate
- Mentioned, without saying whether it is needed
Read out of the employer's own description. Absence means the posting does not say, not that the answer is no.
The role
As the Manager of Model Validation & Verification (VnV) for Behavior Autonomy, you will lead an engineering and data science team responsible for evaluating, benchmarking, and validating the machine learning models and behavioral algorithms that drive our autonomous vehicle (AV) prediction and planning stacks.
You will own the statistical frameworks, offline/online evaluation metrics, and validation pipelines that ensure our behavioral models operate safely, comfortably, and predictably. Operating at the intersection of Data Science, Machine Learning, and Safety Engineering, you will partner closely with Autonomy Software, Prediction, Planner, and ML Operations teams to establish data driven release criteria for our AV fleet.
In this role, you will:
- Team Leadership & Execution: Lead, mentor, and scale a high-performing team of Data Scientists, ML Validation Engineers, and Software Engineers while driving roadmaps, sprint execution, resource allocation, and high-throughput model releases with rigorous safety guardrails. Culture of Rigor: Foster a culture of statistical excellence, healthy skepticism, proactive risk tracking, and data-driven decision-making.
- Validation Strategy & Methodologies: Define and execute end-to-end validation strategies across offline evaluation, open/closed-loop simulation, and shadow-mode fleet benchmarking to ensure robust behavioral model performance. Statistical uncertainties, and regressions into clear, data-driven recommendations for release gating and executive leadership.
- Metrics, Release Gating & Rigor: Oversee metric development and standardization with System Safety and Autonomy teams, establishing quantitative go/no-go release criteria for Behavioral Planner and Prediction ML models while fostering statistical rigor and proactive risk management.
- Cross-Functional & Infrastructure Partnership: Partner closely with Planner, Prediction, MLOps, and Developer Efficiency teams to translate behavioral requirements into measurable validation targets, streamline dataset and evaluation pipelines, and optimize runtime and compute costs.
- Executive Communication & Decision-Making: Translate complex model performance trade-offs, statistical uncertainty, regressions, and safety risks into clear, data-driven recommendations for release decisions and executive leadership.
Qualifications:
- Experience: Masters or PhD in CS, Robotics, Applied Statistics or a related field and 3+ years of direct engineering management experience leading Data Science, Machine Learning, or V&V engineering teams, alongside 7+ years of technical experience in robotics, autonomous systems, or AI/ML.
- Domain Knowledge: Strong background in ML model validation, behavioral evaluation frameworks, system-level performance benchmarking, and statistics.
- Software & Systems Literacy: Strong technical foundation in Python and modern data/ML platforms, with exposure to or conceptual literacy in large-scale production codebases (C++ or distributed systems). Proven ability to partner with systems software engineers, review technical architecture, and understand compute/performance trade-offs, Track record of leading teams evaluating complex robotic systems
- Technical Depth: Proven familiarity with modern C++/Python ML environments, simulation frameworks, high-throughput ML evaluation pipelines.
- Cross-Functional Leadership: Demonstrated ability to navigate complex organizational trade offs between release velocity, compute cost, and safety rigor.
Bonus Qualification:
- Experience with autonomous vehicles, robotics, or other safety-critical systems.
- Experience building large-scale simulation, model evaluation, or validation infrastructure.
- Experience with reinforcement learning, generative AI, or distributed ML systems.
Follow us on LinkedIn
AccommodationsIf you need an accommodation to participate in the application or interview process please reach out to accommodations@zoox.com or your assigned recruiter.
A Final Note:You do not need to match every listed expectation to apply for this position. Here at Zoox, we know that diverse perspectives foster the innovation we need to be successful, and we are committed to building a team that encompasses a variety of backgrounds, experiences, and skills.
Published by Zoox on their own careers page and reproduced here unedited. Read it at Zoox.
Apply at Zoox → Applications go directly to Zoox. This board does not sit in between, take a fee from you, or see your application.
What this listing does not tell you
Listed 25 days. Of the 114 evals and benchmarks roles this board has watched from listing to removal, 32% were gone from their employer's careers page by day 25, and the median came down after 56 days. That is a description of other listings that have already ended, not a prediction about this one: this board records when a listing disappears, never why, and a posting still up is not on a clock it can see.
Pay is not confirmed for this role. Check the employer’s posting for current compensation. Explore published bands from other roles →
Zoox has 2 roles open on this board, 2 of them in evals and benchmarks.
Free. One email on Thursdays.
More roles like this
Matched by discipline, title, listed location and work arrangement.
Same discipline: Evals & Benchmarks · Both have management titles
Santa Clara, CA
Same discipline: Evals & Benchmarks · Both have management titles
Mountain View, CA, USA · San Francisco, CA, USA$251k – $310k
Same discipline: Evals & Benchmarks · Both have management titles
Mountain View, California (HQ)$235k – $352k
Same discipline: Evals & Benchmarks · Both have management titles
Cambridge, MA
Same discipline: Evals & Benchmarks · Both have management titles