Senior AI Engineer – Inference & Agent Systems
Arcana Analytics
Job description
About the role
Arcana Analytics is building real‑time AI agents that synthesize information from heterogeneous sources and deliver structured, reasoned answers. As a Senior AI Engineer you will own the performance, reliability, and correctness of the inference and multi‑step agent pipelines that power this product.
Key responsibilities
- Drive time‑to‑first‑token (TTFT) below 400 ms for multi‑step agent pipelines, including streaming first‑token delivery while sub‑agents continue processing.
- Design and implement KV‑cache strategies, prompt compression, and dynamic context‑window management.
- Implement multi‑provider routing that selects models (OpenAI, Anthropic, Gemini, open‑weight) based on latency, cost, and task type.
- Build Plan‑Execute‑Synthesize DAG pipelines that run sub‑agents in parallel, using Temporal for reliable orchestration (retries, timeouts, idempotency).
- Enforce structured JSON output with schema validation, retry loops, and graceful degradation for malformed LLM responses.
- Own the end‑to‑end evaluation framework: ground‑truth datasets, automated scoring, regression detection, LLM‑as‑judge pipelines, and latency regression testing (p50/p95/p99).
- Develop model‑serving and cold‑start optimizations, async worker architecture for parallel execution, and full observability of token flow, tool calls, and synthesis steps.
Required profile
- Proven experience shipping inference pipelines at production scale with TTFT as a primary metric.
- Hands‑on experience building multi‑step agent systems and troubleshooting their failures in production.
- Track record creating evaluation harnesses and ground‑truth datasets that drive meaningful product improvements.
- Deep understanding of LLM non‑determinism and techniques to make systems resilient to it.
- Experience with streaming LLM responses and handling partial outputs.
Required skills
- Go programming language
- Temporal workflow orchestration
- OpenAI, Anthropic, Gemini APIs (and open‑weight model integration)
- LLM prompt engineering, KV‑cache, prompt compression
- JSON schema validation
- Streaming response handling
- Model serving and cold‑start optimization
- Async worker architecture
- Observability and tracing of token‑level events
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in the United States.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 4 weeks ago
Expires 1 month from now
8 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Arcana Analytics
Related job offers
-
EAM IFS Cloud Consultant, Sr. Associate
pwc United States -
EAM IFS Cloud Consultant, Manager
pwc United States -
Senior Manager – Conversational & Agentic AI
pwc United States -
Staff Software Engineer – React Native Mobile
Candescent Serbie Centrale -
Chief Information Officer (CIO)
St. Olaf College Northfield