Lead Site Reliability Engineer at Zeta Global | AdTechTalent
Engineering2 months agoOn-site
Zeta Global
Lead Site Reliability Engineer
site reliability engineeringSREAWSKubernetesTerraformPulumiOpenTelemetryobservabilityPythonGobashchaos engineeringincident managementCI/CDinfrastructure as codemonitoringmetricscloudLinux
Key details
Salary
Not specified
Employment type
Full-time
Seniority
Lead
Years experience
3-5
Location
Bengaluru, India
Full job description
Lead Site Reliability Engineer responsible for implementing and managing SLOs, SLIs, and error budgets to ensure 99.9%+ uptime. Lead incident response and post-incident reviews with root cause analysis. Automate incident detection and response. Write software to support reliability. Design and implement observability using OpenTelemetry. Perform capacity planning and performance testing. Collaborate with development and operations teams to build reliable and scalable services. Follow best practices in infrastructure design and maintenance using AWS, Kubernetes, EKS, Fargate. Use Infrastructure as Code tools like Terraform or Pulumi. Participate in chaos engineering and on-call rotation. Drive advanced alerting and anomaly detection. Requires 3-5 years SRE experience in cloud and on-prem environments, strong Linux and networking knowledge, experience with AWS, Kubernetes, observability tools (Honeycomb, Grafana, Prometheus, Thanos, ELK, Loki), programming in Python or Go, shell scripting, OpenTelemetry, and chaos engineering tools. Strong skills in incident management, CI/CD, deployment strategies, and resiliency patterns.
What you'll do
Implement and manage SLOs, SLIs, and error budgets to drive reliability
Similar jobs
More roles worth a look
Related opportunities based on specialty and working model so candidates can keep momentum.