GPU Cluster Engineer, Systems & Platform
sciforium · San Francisco
Job description
About the role
Sciforium is building a next‑generation AI infrastructure platform and needs a GPU Cluster Engineer to own the full software stack of its GPU clusters. You will define what a production‑ready node looks like, author images, playbooks and pipelines, and keep the fleet consistent, upgradable and high‑performance for both training and serving workloads.
Key responsibilities
- Design, build and maintain versioned OS images, kernel tuning, GPU/NIC driver stacks and automated bring‑up pipelines that turn a bare server into a validated GPU node.
- Create and run acceptance suites (DCGM, NCCL/RCCL tests, bandwidth, HPL) to gate nodes before they enter the scheduler pool.
- Execute rolling kernel, driver and toolkit upgrades, enforce configuration consistency and manage the CUDA/ROCm‑framework compatibility matrix.
- Develop self‑healing workflows that detect unhealthy nodes (Xid/ECC errors, thermal throttling) and trigger cordon, drain, reboot or re‑image actions.
- Manage node and cluster configuration as code using Ansible or SaltStack with Git‑based review, CI validation and canary rollouts.
- Operate GPU‑enabled Kubernetes for inference (NVIDIA GPU Operator, device plugins, MIG/MPS) and Slurm/Run:AI for multi‑node training, including container integration.
- Maintain the full accelerator stack (CUDA, cuDNN, NCCL, ROCm, RCCL, GPUDirect, MOFED/DOCA) and curated PyTorch/JAX environments.
- Own observability stack (DCGM exporter, Prometheus, Grafana) and troubleshoot complex issues such as NCCL hangs, CUDA memory leaks or ROCm kernel crashes.
Required profile
- 5+ years of systems or infrastructure engineering experience with large‑scale GPU clusters, HPC or ML workloads.
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering or related field.
- Deep Linux internals knowledge (kernel modules, DKMS, systemd, cgroups, NUMA) and hands‑on experience with NVIDIA (CUDA) and/or AMD (ROCm) driver stacks.
- Production experience with Kubernetes GPU workloads and HPC schedulers such as Slurm or Run:AI.
- Strong configuration‑management background using Ansible or SaltStack with Git‑based workflows.
Required skills
- Python, Bash
- Ansible, SaltStack, Git
- Packer, MaaS, Terraform (or similar provisioning tools)
- Docker, containerd, NVIDIA Container Toolkit, ROCm container stack
- Kubernetes, NVIDIA GPU Operator, Slurm, Run:AI
- CUDA, cuDNN, NCCL, ROCm, RCCL, GPUDirect, MOFED/DOCA
- Prometheus, Grafana, DCGM exporter
- InfiniBand, RoCE, RDMA, NVLink, NVSwitch
- PyTorch, JAX
What we offer
- Medical, dental and vision insurance
- 401k plan
- Daily lunch, snacks and beverages
- Flexible time off
- Competitive salary and equity
- Equal opportunity employment
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in the United States.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 19 hours ago
Expires 1 month from now
2 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
sciforium
San Francisco
Related job offers
-
Machine Learning Engineer
sift San Francisco -
Software Engineer – Horizon
sierra San Francisco -
Software Engineer – AI Agent for Tech, Media & Telecom
sierra San Francisco -
Director
Food Safety and Inspection Service Thitani location -
AI Safety Specialist (Short‑Term Project)
HumanitApp Thitani location