AI EvalsJobs
Hiring?
All roles / Crosby / Member of Technical Staff, Research Engineer

Member of Technical Staff, Research Engineer

Evals & Benchmarks3mo ago
Source verified: this exact posting was present on Crosby’s careers feed on . View employer source →

What the posting asks for

Names
Python, veRL, vLLM
Doctorate
Not mentioned

Read out of the employer's own description. Absence means the posting does not say, not that the answer is no.

The role

About Crosby

Crosby is built on a simple conviction: a great legal system is the watermark of a great society — and the best legal work comes from combining human expertise with the right technology, not replacing one with the other.

Crosby is the first AI-native law firm, helping ambitious companies like Cursor, Ramp, and Cognition sign commercial contracts faster.

Legal work is both art and science, and we're mapping the frontier between the two — codifying what can be systematized, amplifying human judgment where it matters most. That's why the right way to bring AI into law isn't to sell software and walk away. It's to own the outcome together: higher-quality work, delivered faster. We build proprietary tools and human-in-the-loop workflows that change what's possible for corporate legal teams.

Crosby was founded by Ryan (Stanford Law, Cooley, startup GC) and John (Penn, first 15 engineers at Ramp). Help us transform one of society's most important industries.

About the role

As a Research Engineer, you'll bridge research and production. You'll turn validated methods, benchmarks, and evaluation infrastructure from our research team into dependable, production-grade features. You’ll be responsible for the entire lifecycle of an AI feature—from building the underlying systems to customizing the models that make them run. Your work will involve both robust system design and the pragmatic application of advanced machine learning techniques, including model fine-tuning when necessary. This role is for a product-minded research engineer who knows how to get the most out of modern AI and is obsessed with shipping reliable, high impact features.

We also publish our work on legal AI at intelligence.crosby.ai — take a look at the kinds of problems you'd be working on.

What you'll do

  • Build core AI Systems: Design, build, and scale our primary AI features, including complex Retrieval Augmented Generation (RAG) pipelines and agentic workflows that make our models measurably better at contract review, playbook retrieval and docx editing.

  • Build RL fine-tuning infrastructure: Stand up and harden RL training infrastructure (PPO, GRPO, and related methods) so researchers can train quickly and reliably at scale.

  • Enable RL in a non-verifiable domain: Build reward pipelines, generators, and formal verification harnesses that make scalable RL possible in a domain like contract review.

  • Customize Models for Impact: Selectively fine tune and adapt language models to improve their performance on highly specific legal tasks where off the shelf solutions fall short

  • Ensure Production Quality: Develop evaluation and monitoring systems that catch regressions and reward hacking before they reach users.

  • Apply AI Pragmatically: Stay on the frontier of AI research and identify the most effective techniques and tools to solve immediate product challenges.

  • Collaborate to Ship: Work closely with product, backend, and our in house legal experts to deliver end to end solutions that make our users more effective.

Who you are

  • Experience: 3+ years of professional software engineering experience building and shipping AI/ML-powered products, with deep proficiency in Python and a proven track record of scalable, production-grade systems.

  • Post-training expertise: Familiarity with post-training techniques such as SFT and GRPO, plus hands-on experience with the systems and frameworks that run them at scale—distributed model training, VERL, and vLLM.

  • Practical AI skills: Hands-on experience with LLM APIs (e.g., OpenAI, Anthropic), RAG, agentic workflows, and the full ML project lifecycle.

  • Analytical and pragmatic: A data-driven approach to problem solving and strong intuition for when to build complex infrastructure versus ship a simple solution.

  • Product-focused: Driven by user impact and skilled at making smart technical trade-offs to deliver value quickly.

  • Ownership: You thrive on taking full responsibility for features, from initial concept through deployment and iteration.

Benefits & Perks

  • Unlimited PTO

  • Lunch & dinner in the office every day

  • Medical insurance (multiple plan options through Anthem & Cigna)

  • HSA or FSA, depending on your plan choice

  • Dental insurance

  • Vision insurance

  • One Medical membership

  • 401(k) match

  • In-office desk setup stipend

Equal Opportunity

Crosby is an equal opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or other applicable legally protected characteristics.

Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the New York City Fair Chance Act.

Pursuant to New York Labor Law Section 194-b, the US Pay Range for this position is listed in the job post. Final compensation will be determined based on skills, experience, and qualifications.

Published by Crosby on their own careers page and reproduced here unedited. Read it at Crosby.

Apply at Crosby → Applications go directly to Crosby. This board does not sit in between, take a fee from you, or see your application.

What this listing does not tell you

Crosby has 2 roles open on this board, 2 of them in evals and benchmarks.

Get the weekly AI Evals Jobs briefNew roles and board updates, with published pay where available. This is the general weekly brief. Or browse them all now.

More roles like this

Matched by discipline, title, listed location and work arrangement.

Same discipline: Evals & Benchmarks · Both have staff / lead titles · Shared listed location

Same discipline: Evals & Benchmarks · Both have staff / lead titles · Shared listed location

Same discipline: Evals & Benchmarks · Both have staff / lead titles

Same discipline: Evals & Benchmarks · Both have staff / lead titles

Same discipline: Evals & Benchmarks · Both have staff / lead titles

Privacy · Terms