Jobiglo

No results.

Member of Technical Staff - GPU Infrastructure

hyperbolic · San Francisco

Senior 🇬🇧 English
IPMI Redfish PXE boot automated OS deployment GPU scheduling Terraform Pulumi CI/CD for infrastructure secrets management observability stack object storage distributed file systems Ceph Weka VAST Data API design CUDA InfiniBand RoCE

Job description

About the role

Hyperbolic Labs is building an open‑access AI cloud that aggregates idle GPU resources worldwide. We need a senior infrastructure engineer to design and operate the core provisioning and virtualization layer that turns raw GPUs into a multi‑tenant, programmable pool for thousands of AI developers.

Key responsibilities

  • Design and implement a multi‑tenancy provisioning system for diverse GPU hardware.
  • Develop automated bare‑metal lifecycle workflows using IPMI/Redfish, PXE boot, and cloud‑init.
  • Build GPU scheduling and orchestration logic that accounts for type, memory, topology, and fragmentation.
  • Create and maintain infrastructure as code with Terraform or Pulumi, including CI/CD pipelines, secrets, and observability.
  • Integrate storage solutions (object storage, high‑IOPS block, distributed file systems) for AI/ML training data and checkpoints.
  • Collaborate with hardware vendors to troubleshoot and optimise hardware integrations.

Required profile

  • Deep expertise in bare‑metal provisioning, remote management (IPMI/Redfish), and automated OS deployment.
  • Strong knowledge of GPU architecture, CUDA, and compute optimisation.
  • Proven experience building and scaling cloud infrastructure or distributed systems in production.
  • Excellent communication skills and ability to work with cross‑functional teams.

Required skills

  • IPMI, Redfish, BMC remote management
  • PXE boot, cloud‑init, automated OS deployment
  • GPU scheduling, topology awareness, fragmentation minimisation
  • Terraform or Pulumi
  • CI/CD for infrastructure, secrets management, observability stack
  • Object storage, high‑IOPS block storage, distributed file systems (e.g., Ceph, Weka, VAST Data)
  • API design
  • CUDA and GPU compute optimisation
  • InfiniBand, RoCE (RDMA) networking

What we offer

  • Opportunity to shape a pioneering AI cloud platform.
  • Collaborative, open‑source‑focused engineering culture.
  • Competitive compensation and equity.

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec hyperbolic.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.
Source : ats:ashby

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

Apply now →

By continuing, you accept our terms of use.

Already have an account? Login

A question about this job?

Ask it here: you will get the full job summary by e-mail, right away.

💬 Chat with us on Telegram

Published 2 weeks ago

Expires 1 month from now

10 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

hyperbolic

San Francisco