AI EvalsJobs
Hiring?
All roles / Anyone AI / Machine Learning Engineer – ML Evaluation & Experiment Design

Machine Learning Engineer – ML Evaluation & Experiment Design

Evals & BenchmarksRemote3w ago
Source verified: this exact posting was present on Anyone AI’s careers feed on . View employer source →

The role

Anyone AI is recruiting experienced Machine Learning Engineers for a specialized project focused on reviewing and evaluating machine learning challenges used in AI model training and evaluation.

The work involves analyzing ML experiments, datasets, metrics, and pipelines to determine whether challenges are technically sound, reproducible, appropriately difficult, and genuinely require strong machine learning reasoning.

What You’ll Work On

You’ll review ML challenges involving:

  • Experiment design and model selection

  • Small and synthetic datasets

  • Data quality and preprocessing

  • Distribution shift and data contamination

  • Label noise and feature leakage

  • Model evaluation and metric selection

  • Hyperparameter tuning

  • Train / validation / test methodology

  • Reproducibility and deterministic pipelines

  • Statistical significance of model improvements

A key part of the role is determining whether a challenge actually rewards good ML reasoning, rather than simply being solvable through brute-force model selection or large hyperparameter searches.

What We’re Looking For

  • 3+ years of hands-on applied machine learning experience

  • Strong experience with:

  • ML experiment design

  • Model selection

  • Hyperparameter tuning

  • Model evaluation

  • Data preprocessing and validation

  • Strong understanding of train, validation, and test splits

  • Ability to identify:

  • Data leakage

  • Label noise

  • Distribution shift

  • Spurious correlations

  • Feature leakage

  • Data contamination

  • Experience evaluating whether performance improvements are statistically meaningful rather than random fluctuations

  • Strong understanding of ML evaluation metrics and when different metrics are appropriate

  • Experience debugging ML workloads across CPU and GPU environments

  • Ability to analyze technical problems and provide clear written feedback

  • Nice to Have

    • Experience creating or participating in Kaggle, DrivenData, or similar ML competitions

    • Experience designing benchmark datasets or ML challenges

    • Background in data-centric AI or dataset quality

    • Experience with synthetic data generation and validation

    • Familiarity with statistical testing, confidence intervals, and effect sizes

    • Experience with ML evaluation pipelines, RLHF, or AI model evaluation

    • Experience developing ML curricula or technical assessments

    • Understanding of common ML failure modes such as:

    • Shortcut learning

    • Spurious correlations

    • Goodhart’s Law

    • Simpson’s paradox

    • Metric gaming

    What You’ll Be Responsible For

    • Reviewing ML challenges and determining whether they are well designed and technically solvable

    • Evaluating whether datasets contain meaningful and learnable signals

    • Identifying unintended shortcuts or artifacts in synthetic datasets

    • Determining whether tasks require genuine diagnosis of the underlying ML problem

    • Reviewing evaluation metrics and improvement thresholds

    • Detecting metric gaming, data leakage, and evaluation flaws

    • Verifying reproducibility across the complete data → model → evaluation pipeline

    • Assessing whether challenge difficulty is appropriately calibrated

    • Providing clear recommendations for improving, recalibrating, or excluding problematic tasks

    Engagement

    Work Type: Remote
    Engagement: Part-time, project-based consulting
    Focus: Applied machine learning, experiment design, data quality, and model evaluation

    This role is a strong fit for ML engineers who enjoy debugging experiments, understanding why models succeed or fail, identifying problems in datasets and evaluation pipelines, and designing rigorous machine learning experiments.

    Published by Anyone AI on their own careers page and reproduced here unedited. Read it at Anyone AI.

    Apply at Anyone AI → Applications go directly to Anyone AI. This board does not sit in between, take a fee from you, or see your application.

    What this listing does not tell you

    Listed 24 days. Of the 114 evals and benchmarks roles this board has watched from listing to removal, 32% were gone from their employer's careers page by day 24, and the median came down after 56 days. That is a description of other listings that have already ended, not a prediction about this one: this board records when a listing disappears, never why, and a posting still up is not on a clock it can see.

    Anyone AI has 9 roles open on this board, 3 of them in evals and benchmarks.

    Get the weekly AI Evals Jobs briefNew roles and board updates, with published pay where available. This is the general weekly brief. Or browse them all now.

    More roles like this

    Matched by discipline, title, listed location and work arrangement.

    Same discipline: Evals & Benchmarks · Shared listed location · Both list remote work; check location eligibility

    Same discipline: Evals & Benchmarks · Shared listed location · Both list remote work; check location eligibility

    Same discipline: Evals & Benchmarks · Both list remote work; check location eligibility

    Same discipline: Evals & Benchmarks · Both list remote work; check location eligibility

    Same discipline: Evals & Benchmarks · Both list remote work; check location eligibility

    Privacy · Terms