San Mateo$175k – $220k+ equity
Member of Technical Staff, Evals Platform
What the posting asks for
- Names
- PyTorch
- Doctorate
- Not mentioned
Read out of the employer's own description. Absence means the posting does not say, not that the answer is no.
The role
About Us:
Fireworks is the platform for specialized intelligence, enabling companies to build, train, and serve AI models tailored to their own data, workflows, and products. Founded by the team behind PyTorch and backed by AMD, Atreides, Benchmark Capital, Index Ventures, Lightspeed, NVIDIA, Sequoia Capital, and TCV, Fireworks powers production AI with hundreds of state-of-the-art open models across text, image, embedding, audio, and multimodal workloads. Today, Fireworks is a Series D company valued at $17.5 billion, bringing together an ambitious, collaborative team that's building the future of enterprise AI.
We are seeking a Member of Technical Staff, Evals & Post-Training Product to help define how developers improve models on Fireworks. This role sits at a unique intersection of scalable system design, deep data science, and model quality.
You will build the infrastructure and workflows that connect evaluation and post-training. Our evaluation setup has grown past its original scope, and we need someone who can take it to the next stage, improving programmatic access and scale. You will work across backend systems, sandbox infrastructure, and user-facing surfaces to make it easier to author evals, understand results, and iterate quickly.
Key Responsibilities
Scale Eval Infrastructure: Take ownership of our internal eval setup and evolve it for the future. You will design systems to eliminate single-host coordination bottlenecks, resolve log-syncing latency, and build seamless programmatic access.
Benchmark Obsession & Reproduction: Act as a "metrics obsessive." Track state-of-the-art (SOTA) benchmarks, read the latest research papers, dig deep into data discrepancies, and insist on rigorously reproducing published results internally.
Pioneer New Benchmarks: Design and build entirely new benchmarks to measure model performance on complex, emerging, or domain-specific use cases.
Own Fine-Tuning Product Experiences: Build and improve user-facing workflows for post-training, including fine-tuning experiences across SFT, RFT, and related model-improvement capabilities.
Work Closely With Users: Partner with customers and internal stakeholders to understand evaluation and fine-tuning needs, triage issues, and convert bespoke workflows into productized, reusable solutions.
Minimum Requirements
Experience: 1–7 years of software engineering or data science experience (We are hiring at multiple levels for this role).
Strong System Design Skills: You know how to architect scalable, programmatic systems and transition legacy setups into robust infrastructure.
Sandbox Infrastructure: Hands-on experience building or working with sandbox environments and sandbox infrastructure for secure code execution and testing.
Analytical & Data Science Mindset: You possess a deep understanding of LLM evaluations, how to design them, and how to use the results to guide model improvement. You are meticulous about data and metrics.
Understanding of the GenAI Lifecycle: You understand the end-to-end workflow—from prompting a base model to curating a dataset, fine-tuning, and productionizing agents.
Preferred Qualifications
Experience: 3+ years of software engineering or applied data science experience.
Frameworks & Orchestration: Experience working with the Harbor framework or similar container registry and orchestration tools.
Public Writing & Analysis: A strong interest in discovering where different models excel and fall short, with a desire to write up and publish these insights publicly (e.g., technical blogs, whitepapers).
Inference & Hardware Knowledge: Interest in the hardware side of AI—understanding GPU constraints, inference optimization techniques, and how they relate to model performance.
Startup DNA: Experience in fast-paced environments where you own features end-to-end.
Why Fireworks?
Solve Hard Problems: Tackle challenges at the forefront of AI infrastructure, from low-latency inference to scalable model serving.
Build What’s Next: Work with bleeding-edge technology that impacts how businesses and developers harness AI globally.
Ownership & Impact: Join a fast-growing, passionate team where your work directly shapes the future of AI—no bureaucracy, just results.
Learn from the Best: Collaborate with world-class engineers and AI researchers who thrive on curiosity and innovation.
Fireworks AI is an equal-opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all innovators.
Published by Fireworks AI on their own careers page and reproduced here unedited. Read it at Fireworks AI.
Apply at Fireworks AI → Applications go directly to Fireworks AI. This board does not sit in between, take a fee from you, or see your application.
What this listing does not tell you
Listed 344 days, which is longer than most. Of the 114 evals and benchmarks roles this board has watched from listing to removal, 99% were gone from their employer's careers page by day 344, and the median came down after 56 days. That is a description of other listings that have already ended, not a prediction about this one: this board records when a listing disappears, never why, and a posting still up is not on a clock it can see.
Fireworks AI has 9 roles open on this board, 6 of them in evals and benchmarks.
Free. One email on Thursdays.
More roles like this
Matched by discipline, title, listed location and work arrangement.
Same discipline: Evals & Benchmarks · Both have staff / lead titles · Shared listed location
San Mateo · New York$160k – $180k
Same discipline: Evals & Benchmarks · Both have staff / lead titles · Shared listed location
RemoteRemote listed$240k – $290k
Same discipline: Evals & Benchmarks · Both have staff / lead titles
San Francisco · Palo Alto$200k – $350k
Same discipline: Evals & Benchmarks · Both have staff / lead titles
London · Europe +1$250k – $535k+ equity
Same discipline: Evals & Benchmarks · Both have staff / lead titles