Engineering Manager, Inference Infrastructure
Anthropic · San Francisco
Job description
About the role
Anthropic is looking for an Engineering Manager to lead the Inference Infrastructure team. This group designs the control plane that decides where AI requests are served, how capacity is allocated, and how caches are placed across Anthropic’s inference fleet. The manager will guide a team of ML‑platform, infrastructure, and distributed‑systems engineers to ensure the request‑to‑model path is efficient, reliable, and adaptable to evolving models, hardware, and cloud environments.
Key responsibilities
- Define and own the technical roadmap for coordinating the inference fleet, including traffic routing, capacity placement, caching strategy, and real‑time demand response.
- Partner with product, inference engine, performance, and capacity teams to translate throughput, latency, utilization, and cost opportunities into shipped improvements with measurable outcomes.
- Establish a quantitative‑modeling culture: require measurable claims before shipping and predict impact of changes.
- Set the technical strategy for the control plane across heterogeneous hardware, multiple cloud providers, and all serving surfaces.
- Run the operational backbone: on‑call rotations, incident response, post‑mortems, and safe deployment practices.
- Provide clarity across the API surface, inference engines, capacity planning, and cloud deployment teams.
- Develop, retain, and grow a high‑performing engineering team; hire at a high technical bar and coach engineers through shifting priorities.
Required profile
- Proven leadership of ML‑platform, infrastructure, or distributed‑systems engineering teams.
- Deep systems expertise to make architectural decisions and anticipate fleet‑wide impact of changes.
- Strong track record of delivering reliable, low‑latency services at scale.
- Ability to collaborate across product, performance, and capacity groups.
- Experience guiding teams through rapid growth and evolving technology stacks.
Required skills
- Distributed systems engineering
- Capacity planning and performance modeling
- Load‑balancing and traffic routing algorithms
- Control‑plane design for large‑scale inference fleets
What we offer
- Opportunity to shape the core infrastructure of a leading AI research organization.
- Collaborative environment with top researchers, engineers, and policy experts.
- Competitive compensation and benefits package.
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in the United States.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 3 weeks ago
Expires 1 month from now
10 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Anthropic
San Francisco