Full job description
Zeta Global seeks a Principal DevOps Engineer to lead the design, build, and operation of scalable CI/CD pipelines and deployment strategies including canary releases, blue/green deployments, and feature flag-driven releases. The role includes managing AWS infrastructure, Kafka event streaming, containerization with Docker, and observability platforms such as Grafana, Prometheus, Loki, and Honeycomb. Responsibilities include defining SLOs/SLIs/SLAs, leading incident response, embedding compliance controls for GDPR, CCPA, and SOC 2, and influencing software architecture and engineering culture. Requires 10+ years of experience in DevOps/SRE with expert knowledge of Kubernetes, Terraform, Docker, Kafka, AWS, and GitLab CI/CD. Benefits include unlimited PTO, medical/dental/vision coverage, employee equity, discounts, wellness classes, and pet insurance. Salary range $180,000 - $210,000. Remote role based in the United States.
What you'll do
- Design, build, and operate production-grade CI/CD pipelines enabling multiple developers to deploy concurrently to production multiple times daily with zero-downtime
- Implement and optimize deployment strategies including canary releases, blue/green deployments, rolling updates, incremental rollouts, and feature flag-gated releases
- Build self-service deployment tooling with safety guardrails, automated rollback triggers, and compliance gates
- Establish deployment observability with real-time canary analysis, automated health scoring, and progressive delivery metrics integrated with Grafana, Prometheus, and Honeycomb
- Champion CI/CD workflows using GitLab CI/CD, Helm charts, and Terraform for version-controlled, auditable, and reproducible deployments
- Define and enforce SLOs/SLIs/SLAs, establish error budgets balancing velocity with reliability
- Lead incident response processes including on-call rotations, runbook development, blameless postmortems, and incident command
- Design and implement observability stacks leveraging Grafana, Prometheus, Loki, and Honeycomb
- Identify and eliminate reliability risks through chaos engineering, load testing, capacity planning, and failure mode analysis
- Reduce operational toil through automation, self-healing infrastructure, and intelligent alerting to minimize MTTD and MTTR
- Manage and optimize AWS infrastructure with Infrastructure as Code best practices
- Design and operate Kafka-based event streaming infrastructure for real-time marketing and analytics workloads
- Ensure robust networking including DNS management, service mesh, load balancing, TCP/IP optimization, routing policies, and VPC architecture
- Manage containerization strategy using Docker including image builds, vulnerability scanning, registry management, and runtime security
- Support data infrastructure operations across Snowflake, MySQL, and other databases
- Embed compliance controls into CI/CD pipelines ensuring GDPR, CCPA, SOC 2 enforcement
- Implement audit trails, change management controls, and deployment approval workflows
- Collaborate with Security and Legal teams on compliance obligations
- Maintain awareness of evolving privacy regulations and adapt infrastructure accordingly
- Serve as a technical leader and DevOps disruptor introducing modern practices to improve developer velocity and operational safety
- Influence software architecture decisions to simplify and streamline operational management
- Communicate technical strategies to engineering leadership and cross-functional teams
- Develop reference architectures, internal standards, and golden path templates
- Participate in on-call rotations and lead incident response
Requirements
- 10+ years of progressive experience in DevOps, SRE, Platform Engineering, or Infrastructure Engineering roles
- Expert-level Kubernetes knowledge including cluster administration, Helm chart authoring, custom controllers/operators, network policies, RBAC, and multi-cluster management on AWS EKS
- Deep expertise in CI/CD pipeline architecture and advanced deployment strategies (canary, blue/green, progressive delivery, feature flag integration) at scale
- Strong proficiency with Infrastructure as Code using Terraform including module design, state management, and multi-environment orchestration
- Expert knowledge of Docker containerization including multi-stage builds, security hardening, image optimization, and container runtime management
- Production experience with Apache Kafka including cluster management, topic design, consumer group strategies, and operational monitoring
- Strong networking fundamentals: DNS (Route 53, internal DNS), TCP/IP, routing, API Gateway, load balancing (ALB/NLB), service mesh, VPC peering, transit gateways, and network troubleshooting
- Extensive AWS experience spanning EKS, EC2, SQS, DynamoDB, IAM, VPC, CloudWatch, and related services in production environments
- Hands-on experience with observability platforms: Grafana, Prometheus, Loki, Honeycomb
- Working familiarity with Node.js, React, Python, Java, and Ruby
- Experience operating within regulated environments with knowledge of GDPR, CCPA, SOC 2, and compliance automation
- Proven ability to influence engineering culture and communicate complex technical strategies clearly
- Experience with GitLab CI/CD pipelines including advanced features such as parent-child pipelines, dynamic environments, and security scanning integration
Tech stack
KubernetesHelmGitLab CI/CDTerraformDockerApache KafkaAWS (EKS, EC2, SQS, DynamoDB, IAM, VPC, CloudWatch)GrafanaPrometheusLokiHoneycombNode.jsReactPythonJavaRubyStatsigIstioLinkerd
Benefits
Unlimited PTOExcellent medical, dental, and vision coverageEmployee EquityEmployee DiscountsVirtual Wellness ClassesPet Insurance