Read the field
Evals Brief
Understand what is changing in AI evaluation—and what you can learn from it. Research, benchmarks, tools, practitioner lessons and opportunities.
Explore the reading desk · New to evals? Start here
Last editorial review: 9 September 2026. This is a curated selection, not a live feed. Source dates are shown below. Company job links do not establish that a vacancy was created by the announcement.
Career explainer · 9 September 2026
Can you contribute to evals without coding?
Some opportunities rely on domain judgment, clear instructions and operational skills. Our new guide uses official employer examples to explain the requirements and distinguish employment from vendor projects.
Explore the routes → · Try an evaluation exercise
Hiring selection · 9 September 2026
Five ways into evaluation and safety
Our first hiring edition explains the teams, work, locations and advertised pay behind five selected vacancies.
Read the edition →
Company announcement · 3 September 2026
ai& and Tenstorrent announce JapanFold
The companies announced a platform serving open-source drug-discovery models on infrastructure in Japan.
Why it matters for this audience: specialist applications raise questions about which tasks and domain-specific outcomes a model should be tested on. That is our editorial interpretation; the announcement does not establish new evaluation vacancies.
Read the original announcement · Current ai& roles on our board
The company's post-training vacancy is a general role; we have not verified a connection to JapanFold.
Foundational reading · 9 January 2026
How teams evaluate AI agents
Anthropic's engineering guide explains agent evaluations, grading and the practical challenges of testing systems that take multiple steps.
Read the engineering guide · Start with our plain-English introduction · Explore evaluation roles
The reading desk
A starting library for understanding the work. These are ongoing reference resources, not announcements of new releases. Documentation can change.
Methods and failure analysis
How tasks, trials, graders and environments fit together—and why a score needs context.
Anthropic: evaluating agents
Evaluation workflows
Understand evaluation datasets and the distinction between development testing and production evaluation.
LangSmith: evaluation concepts
Agent environments
Explore the components of a task, including instructions, an environment and verification.
Harbor: task documentation
Traces and debugging
Use recorded actions to investigate where a system went wrong.
LangSmith: observability
How we cover the field
We cover useful developments even when no company is hiring. Each editorial item should link to its original source, show its date, explain the practical significance and distinguish reported claims from our interpretation. Benchmark results need their task and testing conditions; a higher score alone does not prove a better system for every use.
Tool inclusion is for learning, not an endorsement. Company funding is not revenue or proof of job security. Vacancies are linked to specific announcements only when that connection is verified. Corrections and material updates receive an editorial date.
Learn the fundamentals · Explore companies · Who's hiring this week