Member of Technical Staff, Evaluation Execution

metr· Engineering & Research
Apply Now ↗
📍 BerkeleyEmployee💰 USD 286K–503K

About this role

About METR We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation and misalignment. METR has consistently set precedents for catastrophic AI risk evaluations, including the first independent safety evaluations (working informally with Anthropic and OpenAI in 2022), the first loss-of-control evaluations and first agentic dangerous capability evaluations, the first evaluations using finetuning (mentioned briefly here), the first independent evaluations using internal information about training, the first review partnership for company risk analysis, the first embedded redteaming, and the first evaluations of internal deployments. We’ve been consulted and/or favorably referenced by groups on opposite ends of various spectra, including a16z, Khosla, Gary Marcus, Obama, and Dean Ball, and are known for producing one of the most positive results on AI capabilities (the time horizon trend) and the most negative (our downlift study). We’re generally referenced as the canonical third party assessor, e.g. as the obvious candidate to verify conditional pause agreements, and are trusted with AI incident investigations by frontierlabs and governments.  We believe it is robustly good for policymakers and civil society to have a clear understanding of risks from AI systems, and we are extremely excited to build a team of ambitious, excellent people to tackle one of the most important challenges of our time.   What this role looks like Running models on tasks. Often this means integrating models into our agent scaffolds, running them on our infrastructure and checking the results carefully. (METR both develops our own tasks internally and runs external evaluations.) Communicating results and takeaways. This includes designing useful graphs, writing up conclusions for different audiences (system cards, risk reports, regulators, X, etc), and having great takes on what matters for risk. Building software to improve our evaluations. We don't just try and run the same evaluation over and over again. We also run faster, more informative evaluations over time; this means making the right investments (with the support of our platform team). Project management. Live evaluations require keeping track of a bunch of threads and staying organized. With our recent risk report process, we were running many evaluations at once. Strong and professional communication. We run important and sensitive evaluations, and so the team needs to coordinate with METR leadership, lab contacts, regulators, and others.   Why this role matters As part of informing the world about risk from frontier AI systems, METR often runs and publishes evaluations of frontier models. Our evaluations are a central tool the world uses to understand AI progress. Our Time Horizon methodology has been included in systemcards, called an "obsession" by the NYT, has wide reach online, and is used by governments to inform national policy. We’re expanding the ambition and scale of our evaluations. We have recently begun to measure model propensities and monitorability, and we are increasing the speed, reliability, and quantity of evaluations we aim to do so that we can keep the world informed.   How METR’s evaluations are changing over 2026 Time Horizon is close to saturation, so we’re currently working on Time Horizon 2.0, which we expect to be running on models over the next 6 to 18 months.  We’re gearing up for our first large-scale publication on monitorability, which we believe will be similar to TH in helping folks understand trends over time. We spent the past three months working on a large, industry-wide third-party risk assessment program - which includes us collecting information (and running evaluations!) for both monitorability and propensities/alignment. We expect to do much more work as part of our own risk assessment programs in the future. In general, many ambitious impact stories for METR require us having the capacity to run many more evaluations than we have run historically. For example, while our evaluations currently inform many key decisionmakers about AI capabilities, they are not yet consistently run with the scale, reliability, and speed necessary to play concrete, codified roles in regulatory frameworks. Unlocking this capacity is part of the near-future vision for evaluation execution.

Frequently Asked Questions

What is the salary for the Member of Technical Staff, Evaluation Execution role at metr?
The listed salary for this Member of Technical Staff, Evaluation Execution position at metr is USD 286K–503K. This is an Employee role.
Where is the Member of Technical Staff, Evaluation Execution position at metr located?
This Member of Technical Staff, Evaluation Execution role at metr is based in Berkeley. The position is listed as on-site or hybrid. Check the full job description or apply directly to confirm the work arrangement.
Is the Member of Technical Staff, Evaluation Execution role at metr full-time or part-time?
This is listed as a Employee position. It is posted as a Member of Technical Staff, Evaluation Execution role in the Engineering & Research department at metr.
Which team or department does the Member of Technical Staff, Evaluation Execution at metr belong to?
This Member of Technical Staff, Evaluation Execution position is part of the Engineering & Research department at metr. See the full job description for more information about the team structure and responsibilities.
How do I apply for the Member of Technical Staff, Evaluation Execution position at metr?
Click the "Apply Now" button on this page. You will be redirected to metr's official application portal hosted on lever where you can submit your application directly.
When was the Member of Technical Staff, Evaluation Execution job at metr posted?
This Member of Technical Staff, Evaluation Execution position at metr was posted on Apr 27, 2026. Apply as soon as possible — early applications are often reviewed first.
Member of Technical Staff, Evaluation Execution
metr · 💰 USD 286K–503K
Apply for this role ↗

You'll be redirected to metr's official application page on Lever.