Site Reliability Engineer
seeknow · Atlanta
Job description
About the role
As a Site Reliability Engineer, you'll be responsible for the availability, scalability, and operational health of our AWS‑hosted infrastructure. You'll lead incident response, build automation that lets engineering teams ship safely and often, and drive a culture of measurable reliability across the organization.
Key responsibilities
- Own reliability and performance of production services in AWS, including capacity planning, cost optimization, and architecture reviews.
- Design, build, and maintain fully automated CI/CD pipelines that move code from commit to production with minimal manual intervention.
- Lead incident management: serve on‑call rotation, coordinate outage response, run blameless postmortems, and drive remediation to completion.
- Define and track SLOs, SLIs, and error budgets in partnership with product and engineering teams.
- Build and improve observability through monitoring, logging, alerting, and distributed tracing.
- Manage infrastructure as code and eliminate toil through automation.
- Partner with development teams to embed reliability best practices into system design and release processes.
- Contribute to disaster recovery planning, security hardening, and compliance efforts.
Required profile
- 4+ years experience in SRE, DevOps, or cloud operations supporting production systems.
- Deep hands‑on experience operating workloads in AWS (EC2, ECS/EKS, Lambda, RDS, S3, IAM, VPC, CloudWatch).
- Proven incident management experience: on‑call ownership, incident command, root‑cause analysis, and postmortem processes.
- Demonstrated track record building fully automated CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, CodePipeline, or similar).
- Strong infrastructure‑as‑code skills with Terraform, CloudFormation, or CDK.
- Proficiency in at least one scripting or programming language (Python, Go, Bash).
Required skills
- AWS (EC2, ECS/EKS, Lambda, RDS, S3, IAM, VPC, CloudWatch)
- CI/CD tools (GitHub Actions, GitLab CI, Jenkins, CodePipeline)
- Infrastructure as code (Terraform, CloudFormation, CDK)
- Programming/scripting (Python, Go, Bash)
- Containers and orchestration (Docker, Kubernetes)
- Observability tooling (Datadog, Prometheus, Grafana, ELK stack)
What we offer
- Competitive salary
- Comprehensive health, dental, and vision coverage
- 401(k) with company match
- Flexible PTO and hybrid work arrangement in Atlanta
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in the United States.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 7 hours ago
Expires 1 month from now
3 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
seeknow
Atlanta