Research Engineer – Model Evaluations
Anthropic · Remote-Friendly (Travel-Required) | San Francisco
Job description
About the role
We are looking for a Research Engineer to design, implement, and scale evaluations of Anthropic’s Claude models. You will turn ambiguous notions of intelligence into clear, defensible metrics that guide research, product, and public understanding.
Key responsibilities
- Design and run new evaluations of Claude’s reasoning, agentic behavior, knowledge, and safety properties, and create visualizations for researchers and decision‑makers.
- Build and harden a distributed evaluation execution platform that reliably runs hundreds of evals against live training checkpoints.
- Own dashboards that monitor model health, improve signal‑to‑noise, reduce latency, and make regressions impossible to miss.
- Debug anomalous evaluation results during training runs, determine root causes, and communicate findings under time pressure.
- Improve tooling, libraries, and workflows used by researchers to implement and iterate on evaluations.
Required profile
- Strong Python programming experience, especially in production or research infrastructure.
- Experience building or operating reliable distributed systems, data pipelines, or similar large‑scale infrastructure.
- Clear written and verbal communication skills for explaining technical results to non‑specialists.
- Willingness to be on‑call or provide production support during live training runs.
- Interest in the societal impacts of AI and a commitment to safe, beneficial systems.
Required skills
- Python
- Distributed systems and data pipeline development
- Production infrastructure and on‑call support
- Large language model prompting, sampling, and scaffolding
- Data visualization and dashboard creation
- Evaluation metric design for language models
- Observability, monitoring, and experiment‑tracking
- Statistics and experimental design
- Large‑scale dataset sourcing, curation, and processing
- ML training infrastructure
What we offer
- Annual salary range $500,000 – $850,000 USD.
- Competitive compensation and benefits, including optional equity donation matching, generous vacation and parental leave.
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in the United States.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 8 hours ago
Expires 1 month from now
4 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Anthropic
Remote-Friendly (Travel-Required) | San Francisco
Related job offers
-
Technical Architect – AI Solutions & Training
Anthropic Remote-Friendly (Travel-Required) | San Francisco -
Staff Site Reliability Engineer – Safeguards ML Infrastructure
Anthropic Remote-Friendly (Travel-Required) | San Francisco -
Staff Software Engineer – GTM AI Engineering
Anthropic Remote-Friendly (Travel-Required) | San Francisco -
Senior Integration Developer (Remote)
Doyle Security Services, Inc. (DSS) Îles Vierges des États-Unis -
Integration Developer – Enterprise & Cloud Solutions
Doyle Security Services, Inc. (DSS) Îles Vierges des États-Unis