This job is no longer available
This job expired on 07/09/2026. It no longer accepts applications.
DevOps Engineer – Cloud Operations & Reliability
Vatic Labs · New York
Job description
About the role
We are seeking an experienced DevOps Engineer to own the stability, reliability, and efficiency of our trading and research infrastructure on Google Cloud Platform. This operationally focused role bridges engineering, research, and systems teams, ensuring uninterrupted workflows and proactive incident response.
Key responsibilities
- Manage day‑to‑day health of GCP environment, including resource provisioning, IAM, cost controls, and budget monitoring.
- Serve as the operational liaison between engineering, research, and systems teams to resolve production issues quickly.
- Monitor and maintain production trading and research systems with a focus on uptime and reliability.
- Oversee the research pipeline on GCP, handling job scheduling, resource allocation, and platform stability.
- Administer CI/CD pipelines and Kubernetes clusters for repeatable deployments.
- Manage fleets of VMs, including metrics collection, alerting, and log ingestion policies.
- Maintain high‑performance database instances and ensure data pipeline reliability.
- Benchmark and tune critical services for optimal performance across internal and external trading venues.
- Build and maintain operational tooling for monitoring, backup, deployment automation, and testing.
Required profile
- 3+ years of experience in DevOps, platform operations, or SRE.
- Proven ability to work cross‑functionally with engineering, research, and business teams.
- Hands‑on experience with GCP (IAM, compute, networking, cost management).
- Experience using Terraform or similar IaC tools.
- Strong knowledge of Kubernetes cluster and application management.
- Track record of implementing cloud budget controls and monitoring spend.
- Excellent Python/Bash scripting skills in a Linux environment.
Required skills
- Google Cloud Platform (GCP)
- Terraform (or equivalent IaC)
- Kubernetes
- Python
- Bash
- Linux administration
- CI/CD pipelines
- VM fleet management
- Monitoring, alerting, and log ingestion
- Database administration
What we offer
- Opportunity to work on high‑impact trading and research infrastructure.
- Collaborative environment across engineering, research, and operations.
- Focus on operational excellence and automation.
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in the United States.
A question about this job?
Ask it here: you will get the full job summary by e-mail, right away.
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Vatic Labs
New York