LLM Inference & GPU Systems Consultant
delan-associates-inc
Job description
About the role
We are looking for an experienced AI Infrastructure Runtime Engineer to design, deploy, and operate a large‑scale on‑prem LLM inference platform. The environment runs on NVIDIA H200 GPU clusters within an OpenShift AI ecosystem, serving self‑hosted open‑source models such as Llama. This position focuses exclusively on inference workloads and does not involve model training or fine‑tuning.
Key responsibilities
- Optimize NVIDIA GPU runtime for token generation, including prefill/decode and KV‑cache management.
- Deploy and maintain inference engines such as vLLM and TensorRT‑LLM.
- Maximize GPU throughput through batching, latency tuning, and workload orchestration with RunAI and Kubernetes.
- Manage the full Hugging Face model lifecycle: onboarding, deployment, monitoring, and retirement.
- Operate and support the OpenShift AI container platform for GenAI workloads.
Required profile
- 8+ years of experience as an LLM Systems Engineer or AI Infrastructure Runtime Engineer.
- Extensive hands‑on work with NVIDIA H200 clusters and runtime optimization techniques.
- Proven expertise in OpenShift AI and GPU orchestration tools like RunAI.
- Demonstrated success managing Hugging Face model deployments.
Required skills
- NVIDIA H200 GPU clusters
- GPU runtime optimization (prefill/decode, KV‑cache)
- vLLM inference framework
- TensorRT‑LLM
- RunAI workload orchestration
- Kubernetes GPU scheduling
- OpenShift AI platform
- Hugging Face model lifecycle management
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in the United States.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 3 weeks ago
Expires 1 month from now
4 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
delan-associates-inc
Related job offers
-
EAM IFS Cloud Consultant, Sr. Associate
pwc United States -
EAM IFS Cloud Consultant, Manager
pwc United States -
Senior Manager – Conversational & Agentic AI
pwc United States -
Platform Automation Engineer II – Unified Communications
MAPFRE Webster -
IT Product Engineering Internship – Common Capabilities
MAPFRE Webster