Research Scientist, Interpretability
Anthropic · San Francisco
Job description
About the role
Anthropic is seeking a Research Scientist to join its Interpretability team. You will help reverse‑engineer large language models to build a mechanistic understanding that makes AI systems safer and more trustworthy.
Key responsibilities
- Develop and apply mechanistic interpretability methods to dissect neural network behavior.
- Design and build “circuits” that map model features to algorithmic functions.
- Investigate superposition phenomena and create decompositions for more interpretable components.
- Publish research findings in top venues and contribute to Anthropic’s scientific publications.
- Collaborate with engineers, policy experts, and other researchers to translate insights into safer AI systems.
Required profile
- PhD or equivalent research experience in machine learning, AI, or a related field.
- Demonstrated expertise in interpretability, neural network analysis, or related research.
- Strong analytical and problem‑solving abilities with a track record of publishing high‑impact work.
- Ability to work independently and as part of a fast‑moving research team.
Required skills
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in the United States.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 4 weeks ago
Expires 1 month from now
14 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Anthropic
San Francisco
Related job offers
-
Machine Learning Engineer
sift San Francisco -
Software Engineer – Horizon
sierra San Francisco -
Software Engineer – AI Agent for Tech, Media & Telecom
sierra San Francisco -
Director
Food Safety and Inspection Service Thitani location -
AI Safety Specialist (Short‑Term Project)
HumanitApp Thitani location