AdTechTalent
Engineering16 days agoRemote

Attain

Sr/Staff Site Reliability Engineer, Consumer Apps

site reliability engineeringSREAI agentsautomationTerraformKubernetesIstioGCPAWSDockerPrometheusGrafanaBigQuerySpannerCloudSQLKafkaserverlessCI/CDinfrastructure as codeobservabilitycloud-nativeSOC2PCI compliance

Key details

Salary

Not specified

Employment type

Full-time

Seniority

Senior

Years experience

5-10

Location

Chicago, United States

Full job description

Attain is hiring a Senior/Staff Site Reliability Engineer to build and maintain infrastructure and automation tools supporting Klover's fintech platform. Responsibilities include automating manual processes using AI coding agents, developing Terraform modules, Helm charts, managing Kubernetes clusters with Istio, monitoring GCP databases (BigQuery, Spanner, CloudSQL), and building observability dashboards with Prometheus and Grafana. The role requires 6+ years of experience with cloud-native infrastructure (AWS/GCP), AI agent fluency, Docker, Kubernetes, SQL databases, stream and pub/sub technologies, serverless computing, infrastructure-as-code, observability tools, and compliance processes (SOC2, PCI). The position is full-time, senior level, and hybrid with 4 days in-office in Chicago and 1 remote day, with potential for remote arrangements.

What you'll do

  • Use AI agents as a force multiplier for yourself and others
  • Create, improve, and maintain internal agentic tools and harnesses
  • Add automation to both existing and new systems to eliminate manual processes and toil
  • Write Terraform modules for deploying infrastructure resources via GitLab pipelines
  • Develop Helm charts for deploying services and jobs in Kubernetes cluster
  • Define metrics, network policies, and routing rules for Istio service mesh
  • Monitor and maintain GCP BigQuery, Spanner, and CloudSQL databases
  • Pipe metrics to Google-managed Prometheus and build Grafana dashboards and alerts
  • Experiment with GCP offerings, 3rd party vendors, AI tooling, and open-source projects to automate and secure operations
  • Pair with engineering leads to instrument and monitor critical functionality
  • Participate in architecture design and capacity planning to ensure scalability, maintainability, reliability, and security
  • Build, maintain, and improve CI/CD pipeline

Requirements

  • 6+ years of experience building and maintaining large-scale cloud-native infrastructure (AWS and/or GCP)
  • Fluency directing AI coding agents (e.g. Claude Code, Cursor, or similar) to build, operate, and debug real infrastructure
  • Strong judgment on verification of AI agent work
  • Track record of replacing manual operations with durable automation
  • Experience with Docker, Kubernetes, and Istio or similar service mesh technology
  • Experience with SQL database technologies such as MySQL, Google BigQuery, and Google Spanner
  • Experience with stream technologies such as Kafka and Amazon Kinesis
  • Experience with pub sub technologies such as AWS SNS and Google Pub/Sub
  • Experience with serverless computing technologies such as AWS Lambda and Google Cloud Functions/Google Cloud Run
  • Experience with infrastructure-as-code tools such as Terraform
  • Experience with observability tools such as Datadog, Prometheus, and Grafana
  • Strong computer science and software engineering fundamentals
  • Experience with SOC2 and PCI Compliance processes and requirements

Tech stack

AI coding agentsTerraformGitLab pipelinesHelmKubernetesIstioGCP BigQueryGoogle SpannerCloudSQLPrometheusGrafanaAWSDockerKafkaAmazon KinesisAWS SNSGoogle Pub/SubAWS LambdaGoogle Cloud FunctionsGoogle Cloud RunDatadog

Apply now

Ready to take the next step in your career? Click the button below to continue to the application process.

Similar jobs

More roles worth a look

Related opportunities based on specialty and working model so candidates can keep momentum.