Jobiglo

No results.

Software Engineer, Site Reliability (SRE)

sierra · San Francisco

New
Senior 🇬🇧 English
Terraform AWS IAM VPC Container orchestration Prometheus Grafana Datadog CI/CD tooling Incident management automation LLM infrastructure

Job description

About the role

As a Software Engineer on our Site Reliability team at Sierra, you will be responsible for defining and building the foundation of reliability, observability, and scalability across Sierra’s AI‑driven infrastructure. You’ll partner closely with our core engineering and product teams to ensure our systems are highly available, efficient, and built for growth.

Key responsibilities

  • Own Sierra’s observability stack—monitoring, alerting, logging, and tracing—to give engineers clear visibility into system health and performance.
  • Partner with product and platform engineers to design systems that are reliable and scalable from day one.
  • Design and implement scalable, reliable, and secure cloud infrastructure (AWS) using Terraform and modern DevOps tooling.
  • Improve the reliability and scalability of LLM deployments, ensuring robust, performant, and cost‑effective operation.
  • Lead improvements to deployment pipelines, CI/CD tooling, and incident management processes to reduce downtime and response time.
  • Define the foundation of SRE practices at Sierra, influencing culture, tooling, and best practices across the engineering org.

Required profile

  • 5+ years of hands‑on experience in Site Reliability or Infrastructure engineering for complex SaaS or cloud‑based systems.
  • Experience designing for availability, scalability, and reliability at both infrastructure and application layers.
  • Deep experience with Terraform, AWS services, container orchestration, and cloud networking (IAM, VPC architecture).
  • Strong background in observability systems such as Prometheus, Grafana, Datadog, or similar.
  • Experience working with enterprise customers and understanding compliance and networking needs.
  • Comfortable working in fast‑moving environments and collaborating across product, ML, and core engineering teams.
  • Degree in Computer Science or related field, or equivalent professional experience.

Required skills

  • Terraform
  • AWS (including IAM and VPC)
  • Container orchestration (e.g., Kubernetes)
  • Prometheus
  • Grafana
  • Datadog
  • CI/CD tooling
  • Incident management automation
  • Self‑healing infrastructure patterns
  • LLM infrastructure (inference optimization, model deployment)

What we offer

  • Flexible (unlimited) paid time off
  • Medical, dental, and vision benefits for you and your family
  • Life insurance and disability benefits
  • Retirement plan (depending on country of employment)
  • Parental leave
  • Fertility and family‑building benefits through Carrot
  • Lunch, snacks, and coffee
  • Discretionary benefit stipend
  • Free alphorn lessons

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec sierra.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.
Source : ats:ashby

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 23 hours ago

Expires 1 month from now

1 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

sierra

San Francisco