Jobiglo

No results.

Senior Site Reliability & Automation Engineer (Customer Facing)

Bitdeer Technologies Group · San Jose

Remote
Remote Senior 🇬🇧 English
Kubernetes GPU Nvidia GPU operator Device plugin MIG configuration GPU time-slicing Multi-tenant GPU allocation Topology-aware scheduling NVLink Network rail affinity Quota management Namespaces Network policies RBAC Resource quotas Pod security Bare-Metal-as-a-Service Automation Runbook automation

Job description

About the role

Bitdeer’s NeoCloud is building an AI‑operated GPU cloud where reliability is the product. As a Senior Site Reliability & Automation Engineer you will own the end‑to‑end reliability of the customer‑facing GPU service, ensuring tenants meet their SLAs for availability, job completion, and provisioning latency.

Key responsibilities

  • Design, implement, and operate production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs).
  • Manage Nvidia GPU operator, device plugins, MIG configuration, GPU time‑slicing, and multi‑tenant allocation policies.
  • Implement topology‑aware scheduling, including GPU locality, NVLink domain awareness, and network rail affinity.
  • Handle tenant lifecycle: onboarding, quota management, isolation enforcement (namespaces, network policies, RBAC, resource quotas, pod security), and off‑boarding.
  • Automate Bare‑Metal‑as‑a‑Service provisioning, handoff, lifecycle, and reclamation.
  • Define and monitor SLIs/SLOs/SLAs for cluster availability, job completion rates, provisioning latency, and API availability.
  • Lead incident management with customer communication, runbook automation, and escalation processes.

Required profile

  • Senior‑level SRE or automation engineer with proven experience in large‑scale GPU cloud environments.
  • Strong background in Kubernetes administration and production operations.
  • Hands‑on experience with Nvidia GPU tooling (operator, MIG, device plugins).
  • Demonstrated ability to design observability, automation, and reliability frameworks for multi‑tenant services.

Required skills

  • Kubernetes
  • GPU
  • Nvidia GPU operator
  • Device plugin
  • MIG configuration
  • GPU time‑slicing
  • Multi‑tenant GPU allocation
  • Topology‑aware scheduling
  • NVLink
  • Network rail affinity
  • Quota management
  • Namespaces
  • Network policies
  • RBAC
  • Resource quotas
  • Pod security
  • Bare‑Metal‑as‑a‑Service
  • Automation
  • Runbook automation

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Bitdeer Technologies Group.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 4 weeks ago

Expires 4 weeks from now

24 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Bitdeer Technologies Group

San Jose