AdTechTalent
Engineering16 days agoOn-site

Zeta Global

Lead Site Reliability Engineer

site reliability engineeringSREAWSKubernetesTerraformPulumiOpenTelemetryobservabilityPythonGobashchaos engineeringincident managementCI/CDinfrastructure as codemonitoringmetricscloudLinux

Key details

Salary

Not specified

Employment type

Full-time

Seniority

Lead

Years experience

3-5

Location

Bengaluru, India

Full job description

Lead Site Reliability Engineer responsible for implementing and managing SLOs, SLIs, and error budgets to ensure 99.9%+ uptime. Lead incident response and post-incident reviews with root cause analysis. Automate incident detection and response. Write software to support reliability. Design and implement observability using OpenTelemetry. Perform capacity planning and performance testing. Collaborate with development and operations teams to build reliable and scalable services. Follow best practices in infrastructure design and maintenance using AWS, Kubernetes, EKS, Fargate. Use Infrastructure as Code tools like Terraform or Pulumi. Participate in chaos engineering and on-call rotation. Drive advanced alerting and anomaly detection. Requires 3-5 years SRE experience in cloud and on-prem environments, strong Linux and networking knowledge, experience with AWS, Kubernetes, observability tools (Honeycomb, Grafana, Prometheus, Thanos, ELK, Loki), programming in Python or Go, shell scripting, OpenTelemetry, and chaos engineering tools. Strong skills in incident management, CI/CD, deployment strategies, and resiliency patterns.

What you'll do

  • Implement and manage SLOs, SLIs, and error budgets to drive reliability
  • Develop resilient systems ensuring 99.9%+ uptime for critical services
  • Lead incident response and blameless post-incident reviews with root cause analysis
  • Automate incident detection and response using automated runbooks or workflows
  • Write software to support reliability or efficiency needs
  • Design and implement full observability using tools like OpenTelemetry
  • Use capacity planning, forecasting, and performance testing for scalability
  • Collaborate with development and operations teams to build reliable, scalable services
  • Ensure best practices in infrastructure design, deployment, and maintenance
  • Champion Infrastructure as Code using Terraform, Pulumi, or similar tools
  • Participate in chaos engineering initiatives
  • Participate in on-call rotation
  • Drive advanced alerting and anomaly detection applied to metrics

Requirements

  • 3-5 years of experience as an SRE, working in cloud-based and on-prem environments
  • Deep understanding of Linux systems, networking, and systems administration
  • Experience with cloud platforms like AWS and Kubernetes container orchestration
  • Hands-on experience with observability tools such as Honeycomb, Grafana, Prometheus, Thanos, ELK, or Loki
  • Strong skills in at least one programming language (Python, Go) for production code
  • Strong skills in shell scripting using bash or similar
  • Experience with OpenTelemetry or other distributed tracing systems
  • Experience with Chaos Engineering methodologies and tools
  • Reliability-focused mindset balancing fast product iterations and system stability
  • Solid understanding of SLOs, SLIs, and error budgets
  • Hands-on knowledge of CI/CD pipelines and infrastructure automation
  • Proven expertise in incident management, postmortems, and root cause analysis
  • Knowledge of modern deployment strategies and resiliency patterns

Tech stack

AWSKubernetesEKSFargateTerraformPulumiOpenTelemetryHoneycombGrafanaPrometheusThanosELKLokiPythonGobashChaos MeshChaos MonkeyAWS Fault Injection Simulator

Apply now

Ready to take the next step in your career? Click the button below to continue to the application process.

Similar jobs

More roles worth a look

Related opportunities based on specialty and working model so candidates can keep momentum.