Performance Engineer (Inference, Training & GPU)
World Labs · San Francisco
Job description
About the role
World Labs is seeking a Performance Engineer to accelerate the training and inference of its large generative world models. You will work directly with researchers to identify and eliminate bottlenecks across the entire stack, from low‑level tensor kernels to fleet‑wide serving pipelines.
Key responsibilities
- Optimize end‑to‑end inference and serving for latency, throughput, batching, caching, and scheduling at production scale.
- Write, tune, and fuse GPU kernels using CUDA and Triton, targeting memory‑bound, bandwidth‑bound, and low‑precision (FP8/INT8) execution.
- Improve training throughput and GPU utilization through parallelism strategies, compute/communication overlap, mixed‑precision techniques, and pipeline stall elimination.
- Develop performance models, profiling workflows, and observability tools to make trade‑offs between throughput, latency, cost, and utilization transparent.
- Ensure numerical correctness across precision, kernel, and hardware changes.
- Collaborate with researchers to productionize models and accelerate experimental cycles.
- Contribute to distributed‑system components that support large‑scale training and inference when needed.
Required profile
- Strong foundations in performance engineering: profiling, roofline analysis, latency/throughput optimization, and disciplined root‑cause investigation.
- Deep experience with GPU programming and optimization, especially for inference and training workloads.
- Familiarity with distributed‑system concepts is a plus.
Required skills
- CUDA programming
- Triton kernel development
- GPU kernel optimization
- Low‑precision arithmetic (FP8, INT8)
- Profiling and performance analysis tools
- Roofline analysis
- Mixed‑precision training techniques
- Parallelism and compute/communication overlap strategies
What we offer
- Opportunity to work on cutting‑edge AI models at the frontier of spatial intelligence.
- Collaboration with world‑class researchers and engineers.
- Access to state‑of‑the‑art hardware and resources.
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in the United States.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 1 month ago
Expires 4 weeks from now
13 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
World Labs
San Francisco
Related job offers
-
Senior Full Stack Engineer (Back-End) – AI Platform
Employeur non precise San Francisco -
Senior Product Engineer – Autonomous AI Assistant (On‑Site)
Employeur non precise San Francisco -
Senior/Staff Backend Engineer – Real‑Money Sports Picks
sleeper San Francisco -
Senior Software Engineer - Data Platform
Motional Singapour -
Expert AI Engineer
Helpware