Jobiglo

No results.

ML Infrastructure Engineer

clera · San Mateo

New
Onsite Senior 🇬🇧 English
TensorFlow Serving TorchServe KServe Docker Kubernetes AWS GCP Azure Prometheus Grafana Distributed tracing Python Go Rust C++ Java Knowledge graphs Semantic search Graph databases

Job description

About the role

This is a hands‑on infrastructure engineering role at an early‑stage enterprise AI company building a context and data‑governance layer for AI agents deployed in highly regulated industries. You will own the inference and model‑serving infrastructure end‑to‑end, making production AI agents fast, reliable, and scalable as concurrency grows.

Key responsibilities

  • Design, build, and own inference and model‑serving infrastructure from initial architecture through production deployment.
  • Scale systems that enable AI agents to run reliably and efficiently under increasing concurrent load.
  • Identify and resolve infrastructure bottlenecks in collaboration with ML and platform engineering teams.
  • Drive performance optimization across latency, throughput, and reliability for production workloads.

Required profile

  • 5+ years building and operating ML inference systems, model‑serving platforms, or ML infrastructure in production.
  • Hands‑on experience with frameworks such as TensorFlow Serving, TorchServe, Triton, KServe, or custom solutions.
  • Strong distributed‑systems fundamentals, including managing concurrent requests and resource allocation under load.
  • Proficiency with containerization and orchestration technologies, especially Docker and Kubernetes, for ML workloads.
  • Experience with cloud platforms (AWS, GCP, or Azure) for deploying and managing ML systems.
  • Solid monitoring and observability skills using tools such as Prometheus, Grafana, ELK, or distributed tracing.
  • Proficiency in at least one backend language: Python, Go, Rust, C++, or Java.
  • Familiarity with knowledge graphs, semantic search, or graph databases is a plus.
  • Background in agentic or autonomous AI systems, real‑time inference, or enterprise data infrastructure is a plus.

Required skills

  • TensorFlow Serving
  • TorchServe
  • Triton Inference Server
  • KServe
  • Docker
  • Kubernetes
  • AWS
  • GCP
  • Azure
  • Prometheus
  • Grafana
  • ELK stack
  • Distributed tracing
  • Python
  • Go
  • Rust
  • C++
  • Java
  • Knowledge graphs
  • Semantic search
  • Graph databases

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec clera.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.
Source : ats:ashby

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 6 hours ago

Expires 1 month from now

2 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

clera

San Mateo