AI EvalsJobs
Hiring?
All roles / Anyone AI / Senior Software Engineer – Open Source & SWE-Bench Evaluation

Senior Software Engineer – Open Source & SWE-Bench Evaluation

Evals & BenchmarksRemote3w ago
Source verified: this exact posting was present on Anyone AI’s careers feed on . View employer source →

What the posting asks for

Names
Python
Doctorate
Not mentioned

Read out of the employer's own description. Absence means the posting does not say, not that the answer is no.

The role

Anyone AI is recruiting experienced Software Engineers for a specialized project focused on reviewing and evaluating real-world software engineering tasks derived from open-source repositories.

The work involves assessing whether coding tasks based on real GitHub issues and pull requests are technically sound, reproducible, appropriately tested, and representative of the kinds of problems professional software engineers solve every day.

What You’ll Work On

You’ll review software engineering tasks involving:

  • Real-world bug fixes and feature implementations

  • Open-source repositories and pull requests

  • Unit tests and test coverage

  • Repository setup and dependency management

  • Reproducibility and environment configuration

  • Task difficulty and complexity

  • Multi-file and cross-module code changes

  • Technical feedback and quality assessment

You’ll determine whether tasks are clearly specified, technically solvable, supported by sufficient tests, and free from issues such as flaky tests, missing dependencies, ambiguous requirements, or environment-specific behavior.

What We’re Looking For

  • 3+ years of professional software engineering experience

  • Strong experience working with large, multi-file codebases

  • Experience reviewing pull requests, debugging issues, and maintaining production code

  • Strong understanding of unit testing and test coverage

  • Ability to evaluate whether tests correctly validate a solution without unnecessarily restricting implementation approaches

  • Experience with dependency management, environment setup, and reproducibility

  • Strong understanding of Git and GitHub-based development workflows

  • Ability to analyze complex technical problems and provide clear written feedback

Nice to Have

  • Contributions to or maintenance of open-source projects

  • Experience with SWE-Bench, SWE-Bench Verified, or similar coding benchmarks

  • Experience with major Python open-source projects such as Django, Flask, scikit-learn, SymPy, matplotlib, requests, or pytest

  • Experience with Docker, CI/CD, pip, conda, or dependency pinning

  • Knowledge of test fixtures, test isolation, or property-based testing

  • Experience designing technical assessments or reviewing coding challenges

  • Experience with AI/ML evaluation, data curation, RLHF, or benchmark development

What You’ll Be Responsible For

  • Reviewing coding tasks derived from real GitHub issues and pull requests

  • Assessing whether problem statements and success criteria are clear and complete

  • Evaluating unit tests for correctness, coverage, and robustness

  • Identifying flaky tests, missing dependencies, version conflicts, and environment issues

  • Determining whether tasks can be reliably reproduced across environments

  • Assessing the real-world difficulty and complexity of each task

  • Providing clear recommendations on whether tasks should be accepted, improved, or excluded

Engagement

Work Type: Remote
Engagement: Part-time, project-based consulting
Focus: Software engineering, open-source code review, testing, and technical evaluation

This role is a strong fit for experienced engineers who enjoy debugging complex codebases, reviewing pull requests, working with open-source software, and evaluating what makes a software engineering problem well designed.

Published by Anyone AI on their own careers page and reproduced here unedited. Read it at Anyone AI.

Apply at Anyone AI → Applications go directly to Anyone AI. This board does not sit in between, take a fee from you, or see your application.

What this listing does not tell you

Listed 24 days. Of the 114 evals and benchmarks roles this board has watched from listing to removal, 32% were gone from their employer's careers page by day 24, and the median came down after 56 days. That is a description of other listings that have already ended, not a prediction about this one: this board records when a listing disappears, never why, and a posting still up is not on a clock it can see.

Anyone AI has 9 roles open on this board, 3 of them in evals and benchmarks.

Get the weekly AI Evals Jobs briefNew roles and board updates, with published pay where available. This is the general weekly brief. Or browse them all now.

More roles like this

Matched by discipline, title, listed location and work arrangement.

Same discipline: Evals & Benchmarks · Both have senior titles · Both list remote work; check location eligibility

Same discipline: Evals & Benchmarks · Shared listed location · Both list remote work; check location eligibility

Same discipline: Evals & Benchmarks · Both have senior titles · Both list remote work; check location eligibility

Same discipline: Evals & Benchmarks · Both have senior titles

Same discipline: Evals & Benchmarks · Both have senior titles

Privacy · Terms