Jobiglo

No results.

Senior AI Engineer – Inference & Agent Systems

Arcana Analytics

Senior 🇬🇧 English
Go Temporal OpenAI Anthropic Gemini LLM JSON schema Streaming KV cache Prompt compression Model serving Async worker architecture Observability

Job description

About the role

Arcana Analytics is building real‑time AI agents that synthesize information from heterogeneous sources and deliver structured, reasoned answers. As a Senior AI Engineer you will own the performance, reliability, and correctness of the inference and multi‑step agent pipelines that power this product.

Key responsibilities

  • Drive time‑to‑first‑token (TTFT) below 400 ms for multi‑step agent pipelines, including streaming first‑token delivery while sub‑agents continue processing.
  • Design and implement KV‑cache strategies, prompt compression, and dynamic context‑window management.
  • Implement multi‑provider routing that selects models (OpenAI, Anthropic, Gemini, open‑weight) based on latency, cost, and task type.
  • Build Plan‑Execute‑Synthesize DAG pipelines that run sub‑agents in parallel, using Temporal for reliable orchestration (retries, timeouts, idempotency).
  • Enforce structured JSON output with schema validation, retry loops, and graceful degradation for malformed LLM responses.
  • Own the end‑to‑end evaluation framework: ground‑truth datasets, automated scoring, regression detection, LLM‑as‑judge pipelines, and latency regression testing (p50/p95/p99).
  • Develop model‑serving and cold‑start optimizations, async worker architecture for parallel execution, and full observability of token flow, tool calls, and synthesis steps.

Required profile

  • Proven experience shipping inference pipelines at production scale with TTFT as a primary metric.
  • Hands‑on experience building multi‑step agent systems and troubleshooting their failures in production.
  • Track record creating evaluation harnesses and ground‑truth datasets that drive meaningful product improvements.
  • Deep understanding of LLM non‑determinism and techniques to make systems resilient to it.
  • Experience with streaming LLM responses and handling partial outputs.

Required skills

  • Go programming language
  • Temporal workflow orchestration
  • OpenAI, Anthropic, Gemini APIs (and open‑weight model integration)
  • LLM prompt engineering, KV‑cache, prompt compression
  • JSON schema validation
  • Streaming response handling
  • Model serving and cold‑start optimization
  • Async worker architecture
  • Observability and tracing of token‑level events

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Arcana Analytics.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.
Source : ats:greenhouse

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 4 weeks ago

Expires 1 month from now

8 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Arcana Analytics