Staff+ Software Engineer, Node Infra
Anthropic · San Francisco
Job description
About the role
Anthropic’s Infrastructure organization builds the foundation for safe, reliable AI systems. The Node Infra team owns the full lifecycle of accelerator capacity, provisioning compute across clouds and custom datacenters, and ensuring every GPU, TPU, and Trainium node is ready for frontier AI research.
Key responsibilities
- Define and execute the technical strategy and roadmap for node lifecycle management, including ingestion, bring‑up, health checking, and automated repair.
- Lead cross‑team initiatives to design, scale, and operate AI clusters across multiple cloud providers and accelerator families.
- Build systems that automatically detect, isolate, and remediate unhealthy hardware, improving fleet MTBI and reducing stranded capacity.
- Shape infrastructure architecture, solving the hardest problems either directly or through collaboration.
- Partner with cloud providers and internal research, inference, and product teams to influence long‑term compute and data strategy.
- Establish operational‑excellence practices such as incident response, post‑mortems, and on‑call rotation.
- Mentor and coach engineers, fostering technical growth within the team.
Required profile
- Deep expertise in distributed systems, reliability engineering, and cloud platforms (Kubernetes, IaC, AWS/GCP/Azure).
- Strong proficiency in at least one systems language (Rust, Go, or Python) and IaC with Terraform.
- Hands‑on experience with machine‑learning accelerators (GPUs, TPUs, or Trainium).
- Proven track record leading multi‑quarter technical initiatives across multiple teams.
- Ability to build alignment with senior stakeholders and communicate effectively at all levels.
- 10+ years of software engineering experience, including technical leadership.
- Experience managing hyperscale compute infrastructure (10K+ nodes) and driving capacity efficiency.
- Deep knowledge of Kubernetes internals such as scheduler, autoscaler, or custom controllers.
Required skills
- Kubernetes
- Infrastructure as Code (IaC)
- AWS
- GCP
- Azure
- Rust
- Go
- Python
- Terraform
- GPU
- TPU
- Trainium
- Distributed systems
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in the United States.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 3 weeks ago
Expires 1 month from now
9 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Anthropic
San Francisco