Jobiglo

No results.

GPU Cluster Engineer, Systems & Platform

sciforium · San Francisco

New
Senior 🇬🇧 English
Python Bash Ansible SaltStack Git Packer MaaS Terraform Docker containerd NVIDIA Container Toolkit ROCm container stack Kubernetes NVIDIA GPU Operator Slurm Run:AI CUDA cuDNN NCCL ROCm RCCL GPUDirect MOFED/DOCA Prometheus Grafana DCGM exporter InfiniBand RoCE RDMA NVLink NVSwitch PyTorch JAX

Job description

About the role

Sciforium is building a next‑generation AI infrastructure platform and needs a GPU Cluster Engineer to own the full software stack of its GPU clusters. You will define what a production‑ready node looks like, author images, playbooks and pipelines, and keep the fleet consistent, upgradable and high‑performance for both training and serving workloads.

Key responsibilities

  • Design, build and maintain versioned OS images, kernel tuning, GPU/NIC driver stacks and automated bring‑up pipelines that turn a bare server into a validated GPU node.
  • Create and run acceptance suites (DCGM, NCCL/RCCL tests, bandwidth, HPL) to gate nodes before they enter the scheduler pool.
  • Execute rolling kernel, driver and toolkit upgrades, enforce configuration consistency and manage the CUDA/ROCm‑framework compatibility matrix.
  • Develop self‑healing workflows that detect unhealthy nodes (Xid/ECC errors, thermal throttling) and trigger cordon, drain, reboot or re‑image actions.
  • Manage node and cluster configuration as code using Ansible or SaltStack with Git‑based review, CI validation and canary rollouts.
  • Operate GPU‑enabled Kubernetes for inference (NVIDIA GPU Operator, device plugins, MIG/MPS) and Slurm/Run:AI for multi‑node training, including container integration.
  • Maintain the full accelerator stack (CUDA, cuDNN, NCCL, ROCm, RCCL, GPUDirect, MOFED/DOCA) and curated PyTorch/JAX environments.
  • Own observability stack (DCGM exporter, Prometheus, Grafana) and troubleshoot complex issues such as NCCL hangs, CUDA memory leaks or ROCm kernel crashes.

Required profile

  • 5+ years of systems or infrastructure engineering experience with large‑scale GPU clusters, HPC or ML workloads.
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering or related field.
  • Deep Linux internals knowledge (kernel modules, DKMS, systemd, cgroups, NUMA) and hands‑on experience with NVIDIA (CUDA) and/or AMD (ROCm) driver stacks.
  • Production experience with Kubernetes GPU workloads and HPC schedulers such as Slurm or Run:AI.
  • Strong configuration‑management background using Ansible or SaltStack with Git‑based workflows.

Required skills

  • Python, Bash
  • Ansible, SaltStack, Git
  • Packer, MaaS, Terraform (or similar provisioning tools)
  • Docker, containerd, NVIDIA Container Toolkit, ROCm container stack
  • Kubernetes, NVIDIA GPU Operator, Slurm, Run:AI
  • CUDA, cuDNN, NCCL, ROCm, RCCL, GPUDirect, MOFED/DOCA
  • Prometheus, Grafana, DCGM exporter
  • InfiniBand, RoCE, RDMA, NVLink, NVSwitch
  • PyTorch, JAX

What we offer

  • Medical, dental and vision insurance
  • 401k plan
  • Daily lunch, snacks and beverages
  • Flexible time off
  • Competitive salary and equity
  • Equal opportunity employment

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec sciforium.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.
Source : ats:ashby

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 15 hours ago

Expires 1 month from now

1 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

sciforium

San Francisco