Jobiglo

No results.

Staff+ Software Engineer, Node Infra

Anthropic · San Francisco

Senior 🇬🇧 English
Kubernetes AWS GCP Azure Rust Go Python Terraform GPU TPU Trainium Distributed systems

Job description

About the role

Anthropic’s Infrastructure organization builds the foundation for safe, reliable AI systems. The Node Infra team owns the full lifecycle of accelerator capacity, provisioning compute across clouds and custom datacenters, and ensuring every GPU, TPU, and Trainium node is ready for frontier AI research.

Key responsibilities

  • Define and execute the technical strategy and roadmap for node lifecycle management, including ingestion, bring‑up, health checking, and automated repair.
  • Lead cross‑team initiatives to design, scale, and operate AI clusters across multiple cloud providers and accelerator families.
  • Build systems that automatically detect, isolate, and remediate unhealthy hardware, improving fleet MTBI and reducing stranded capacity.
  • Shape infrastructure architecture, solving the hardest problems either directly or through collaboration.
  • Partner with cloud providers and internal research, inference, and product teams to influence long‑term compute and data strategy.
  • Establish operational‑excellence practices such as incident response, post‑mortems, and on‑call rotation.
  • Mentor and coach engineers, fostering technical growth within the team.

Required profile

  • Deep expertise in distributed systems, reliability engineering, and cloud platforms (Kubernetes, IaC, AWS/GCP/Azure).
  • Strong proficiency in at least one systems language (Rust, Go, or Python) and IaC with Terraform.
  • Hands‑on experience with machine‑learning accelerators (GPUs, TPUs, or Trainium).
  • Proven track record leading multi‑quarter technical initiatives across multiple teams.
  • Ability to build alignment with senior stakeholders and communicate effectively at all levels.
  • 10+ years of software engineering experience, including technical leadership.
  • Experience managing hyperscale compute infrastructure (10K+ nodes) and driving capacity efficiency.
  • Deep knowledge of Kubernetes internals such as scheduler, autoscaler, or custom controllers.

Required skills

  • Kubernetes
  • Infrastructure as Code (IaC)
  • AWS
  • GCP
  • Azure
  • Rust
  • Go
  • Python
  • Terraform
  • GPU
  • TPU
  • Trainium
  • Distributed systems

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Anthropic.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.
Source : ats:greenhouse

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 3 weeks ago

Expires 1 month from now

9 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Anthropic

San Francisco