AI Infrastructure Engineer – Agents & ML Systems
havocai
Job description
About the role
As an AI Infrastructure Engineer focusing on agents and machine‑learning systems, you will design and operate the internal AI platform that enables Havoc’s AI teams to safely and reliably use modern AI models. You will build production‑grade tools, services, pipelines, and integrations that connect large language models, agentic workflows, data sources, and engineering systems.
Key responsibilities
- Build internal AI infrastructure that links LLMs and AI agents with APIs, data lakes, telemetry stores, simulation tools, code repositories, documentation, logs, and engineering workflows.
- Develop and maintain agentic AI systems for task automation, data analysis, engineering support, simulation workflows, and internal productivity.
- Create tool‑integration and connector frameworks for AI agents, including MCP and emerging tool‑use standards, ensuring secure tool‑use patterns.
- Design pipelines for retrieval, retrieval‑augmented generation (RAG), context management, document processing, embeddings, and internal knowledge search.
- Support ML infrastructure workflows such as data preparation, dataset curation, experiment tracking, model evaluation, fine‑tuning, and model deployment.
- Build evaluation frameworks for agent performance, tool‑use reliability, task success, model quality, regression testing, and failure analysis.
- Develop observability, logging, tracing, auditability, monitoring, and debugging tools for AI agents, model calls, MCP tools, and ML pipelines.
- Collaborate with Autonomy, Software, Data, Simulation, Product, and Operations teams to align infrastructure with mission‑critical needs.
Required profile
- Strong software or infrastructure engineering background with hands‑on experience building production systems.
- Curiosity about the practical application of AI and comfort working across the AI stack.
- Ability to design, implement, and operate scalable services that connect models, tools, data, and users.
- Experience collaborating with cross‑functional teams in fast‑moving, high‑impact environments.
Required skills
- Large language models (LLMs)
- MCP (Model Connectors Protocol) or similar tool‑use standards
- Retrieval‑augmented generation (RAG) pipelines
- Context management and document processing
- Embeddings and knowledge‑search systems
- Experiment tracking and model evaluation
- Fine‑tuning and model deployment workflows
- Observability, logging, tracing, and monitoring tools for AI systems
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in the United States.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 3 weeks ago
Expires 1 month from now
4 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
havocai