Jobiglo

No results.

Member of Technical Staff – Compute Cluster

causal · San Francisco

🇬🇧 English
GPU clusters Kubernetes Slurm Docker Linux Networking Storage Infrastructure-as-code GCP AWS Azure Monitoring Logging Observability CUDA NCCL Performance profiling

Job description

About the role

We are looking for a Member of Technical Staff to lead the design, deployment and operation of our large‑scale GPU compute clusters that power cutting‑edge AI research. The role sits within the infrastructure team at causal in San Francisco and focuses on delivering reliable, high‑performance compute for training and inference workloads.

Key responsibilities

  • Design, provision, image, upgrade and capacity‑plan distributed GPU clusters end‑to‑end.
  • Extend scheduling and orchestration systems such as Kubernetes and Slurm for topology‑aware placement, preemption, quotas and multi‑tenancy.
  • Develop software abstractions and a self‑serve interface for researchers and engineers to access cluster resources.
  • Manage cluster storage, checkpoint and log artifact paths with clear retention and lineage policies.
  • Implement observability, monitoring and automated recovery to improve reliability and catch failures early.
  • Collaborate with research teams to unblock large‑scale runs and advise on performance and placement trade‑offs.

Required profile

  • Proven experience operating large‑scale GPU clusters and container orchestration frameworks.
  • Strong systems background covering Linux, networking, storage and infrastructure‑as‑code.
  • Familiarity with major cloud platforms (GCP, AWS, Azure) and their ML/AI services.
  • Understanding of monitoring, logging, observability and version‑control best practices for machine‑learning systems.
  • Ability to own deliverables from requirements through autonomous execution and thrive in fast‑paced, ambiguous environments.

Required skills

  • GPU clusters
  • Kubernetes
  • Slurm
  • Docker
  • Linux
  • Networking
  • Storage
  • Infrastructure‑as‑code
  • Google Cloud Platform (GCP)
  • Amazon Web Services (AWS)
  • Microsoft Azure
  • Monitoring and logging tools
  • Observability platforms
  • Git
  • CUDA
  • NCCL
  • Performance profiling for distributed workloads

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec causal.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.
Source : ats:ashby

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

Apply now →

By continuing, you accept our terms of use.

Already have an account? Login

A question about this job?

Ask it here: you will get the full job summary by e-mail, right away.

💬 Chat with us on Telegram

Published 1 month ago

Expires 1 week from now

18 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

causal

San Francisco