Lead Software Engineer, Model Serving Platform
sciforium · San Francisco
Job description
About the role
This is a rare chance to help architect and lead the development of Sciforium’s next‑generation model serving platform, the high‑performance engine that will bring a multimodal, highly efficient foundation model to market. As a senior technical leader, you’ll not only build core components yourself but also guide and mentor other engineers, influencing engineering direction, standards, and execution quality.
Key responsibilities
- Lead the technical direction of the model serving platform, owning architecture decisions and guiding engineering execution.
- Build core serving components including execution runtimes, batching, scheduling, and distributed inference systems.
- Develop high‑performance C++ and CUDA/HIP modules, including custom GPU kernels and memory‑optimized runtimes.
- Collaborate with ML researchers to productionize new multimodal models and ensure low‑latency, scalable inference.
- Build Python APIs and services that expose model capabilities to downstream applications.
- Mentor and support other engineers through code reviews, design discussions, and hands‑on technical guidance.
- Drive performance profiling, benchmarking, and observability across the inference stack.
- Ensure high reliability and maintainability through testing, monitoring, and engineering best practices.
Required profile
- Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience.
- 5+ years designing and building scalable, reliable backend systems or distributed infrastructure.
- Strong understanding of LLM inference mechanics (prefill vs decode, batching, KV cache).
- Experience with Kubernetes/Ray and containerization.
- Effective communication skills and ability to lead technical discussions, mentor engineers, and drive engineering quality.
Required skills
- C++
- Python
- CUDA/HIP GPU programming
- Custom GPU kernel development
- Kubernetes
- Ray
- Containerization (Docker, etc.)
- Performance profiling and debugging at system level
- Distributed inference and scheduling
What we offer
- Medical, dental, and vision insurance
- 401k plan
- Daily lunch, snacks, and beverages
- Flexible time off
- Competitive salary and equity
- Equal opportunity employment
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in the United States.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 12 hours ago
Expires 1 month from now
1 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
sciforium
San Francisco
Related job offers
-
Machine Learning Engineer
sift San Francisco -
Software Engineer – Horizon
sierra San Francisco -
Software Engineer – AI Agent for Tech, Media & Telecom
sierra San Francisco -
Director
Food Safety and Inspection Service Thitani location -
AI Safety Specialist (Short‑Term Project)
HumanitApp Thitani location