AI EvalsJobs
Hiring?
All roles / Appen / Applied AI Research Engineer

Applied AI Research Engineer

Evals & BenchmarksRemote1w ago
Source verified: this exact posting was present on Appen’s careers feed on . View employer source →

What the posting asks for

Doctorate
Mentioned, without saying whether it is needed

Read out of the employer's own description. Absence means the posting does not say, not that the answer is no.

The role

About the Role

As an Applied Research Engineer, you’ll build practical AI research assets that support Frontier lab initiatives and customer engagements. This is an implementation-focused role for someone who enjoys turning research concepts into working systems.

You’ll work with a high degree of autonomy, experimenting with new approaches and developing solutions that can be reused across customer opportunities. You’ll partner closely with the GenAI Research team and cross-functional stakeholders to bring technical ideas into practical applications.

Your Impact

  • Build reinforcement learning and agent environments for real customer and Frontier lab use cases, including task specifications, scoring, and evaluation.
  • Develop benchmarks and evaluation harnesses to measure model and data quality across areas such as accuracy, robustness, safety, latency, and cost.
  • Build LLM pipelines and agentic systems that support research, evaluation, and customer trials.
  • Run fine-tuning, adapter, and other model experiments to evaluate how data and methods influence model behavior.
  • Deploy local or self-hosted models for evaluation, inference, and automation workflows.
  • Document experiments, configurations, data, results, and known limitations so other engineers can reproduce and build on your work.
  • Partner with the GenAI Research team and cross-functional stakeholders to turn technical work into reusable assets for customer engagements.

What You Bring

  • Bachelor’s, Master’s, or PhD in Computer Science, Engineering, Machine Learning, or a related technical field.
  • 3+ years of professional engineering or relevant industry experience in AI/ML or software engineering.
  • Strong software engineering skills and experience building reliable, maintainable AI systems.
  • Hands-on experience building agentic systems, reinforcement learning environments, LLM pipelines, or similar AI systems.
  • Experience building evaluation harnesses, benchmarks, or model testing pipelines.
  • Ability to work independently on technical problems and move quickly from an idea or research question to a working solution.
  • Strong understanding of experimentation, reproducibility, and technical documentation.

Nice to Haves

  • Developed synthetic data generation systems or datasets.
  • Published research papers, benchmarks, or other technical research.
  • Worked with SWE-bench or similar software engineering evaluation environments.
  • Built or deployed local inference, open-weight models, or self-hosted model environments.

Why You'll Love Working Here

At Appen, we foster a culture of innovation, collaboration, and excellence. We value curiosity, accountability, and a commitment to delivering the highest quality AI solutions for frontier models.

You'll work on complex challenges that shape the future of AI across industries and geographies, alongside talented people in a culture that values humility over ego. You'll have the flexibility to deliver in a way that works for you and your team, supported by tools, resources, and development opportunities to continue to build your capability over time.

About Appen

Appen has been a leader in AI training data for over 30 years. We specialise in human generated data to train, fine tune, and evaluate models across generative AI, large language models, computer vision, and speech recognition. Our AI assisted data annotation platform and global crowd of more than 1 million contributors in over 200 countries support model pre training, supervised fine tuning, evaluation and benchmarking, safety and red teaming, and multilingual global expansion.

Published by Appen on their own careers page and reproduced here unedited. Read it at Appen.

Apply at Appen → Applications go directly to Appen. This board does not sit in between, take a fee from you, or see your application.

What this listing does not tell you

Listed 7 days, which is recent for this board. Of the 114 evals and benchmarks roles this board has watched from listing to removal, 15% were gone from their employer's careers page by day 7, and the median came down after 56 days. That is a description of other listings that have already ended, not a prediction about this one: this board records when a listing disappears, never why, and a posting still up is not on a clock it can see.

Pay is not confirmed for this role. Check the employer’s posting for current compensation. Explore published bands from other roles →

Appen has 4 roles open on this board.

Get the weekly AI Evals Jobs briefNew roles and board updates, with published pay where available. This is the general weekly brief. Or browse them all now.

More roles like this

Matched by discipline, title, listed location and work arrangement.

Same discipline: Evals & Benchmarks · Both list remote work; check location eligibility

Same discipline: Evals & Benchmarks · Both list remote work; check location eligibility

Same discipline: Evals & Benchmarks · Both list remote work; check location eligibility

Same discipline: Evals & Benchmarks · Both list remote work; check location eligibility

Same discipline: Evals & Benchmarks · Both list remote work; check location eligibility

Privacy · Terms