Full job description
FreeWheel, a Comcast company, seeks an experienced Data Site Reliability Engineer (SRE) to ensure the reliability, scalability, and performance of data systems. Responsibilities include designing monitoring and alerting systems, resolving data pipeline and storage issues, developing automation tools for deployment and recovery, optimizing data storage and query performance, incident response, capacity planning, documentation, security compliance, and cross-team collaboration. Required qualifications include 8+ years in SRE, DevOps, or Data Operations, experience with cloud platforms (AWS, GCP, Azure), big data technologies (Kafka, Hadoop, Spark), distributed storage (Cassandra, HDFS, AWS S3), database management (NoSQL, MySQL, PostgreSQL), automation tools (Ansible, Terraform, Kubernetes, Docker), CI/CD pipelines, programming in Python, Go, Java, or Scala, monitoring tools (Prometheus, Grafana, ELK Stack), strong troubleshooting skills, excellent communication, and a bachelor's degree in a related field. Preferred skills include Aerospike, Snowflake, containerization, microservices, large-scale distributed systems, and data governance.
What you'll do
- Design and implement monitoring and alerting systems for data platforms
- Respond to and resolve issues impacting data pipelines or storage layers
- Develop and maintain automation tools and scripts for deployment, monitoring, backup, recovery, and disaster recovery
- Analyze and optimize performance of data storage, query performance, and data flows
- Reduce latency and improve processing speed
- Respond quickly to data platform failures and perform troubleshooting
- Coordinate cross-team efforts to resolve issues
- Ensure high availability and reliability of data platforms
- Work with data engineering teams on capacity planning and scaling
- Document architecture, configurations, and operational procedures
- Share knowledge and provide training
- Ensure data platforms meet security standards and compliance requirements
- Prevent data breaches, unauthorized access, and data misuse
- Collaborate with data science, product, and development teams
- Support data product design and implementation
- Resolve reliability-related issues across the platform
Requirements
- At least 8+ years of experience as an SRE, DevOps, or Data Operations Engineer
- Experience with cloud platforms (AWS, GCP, Azure)
- Familiarity with modern data architectures and big data platforms (Kafka, Hadoop, Spark)
- Experience with distributed storage (Cassandra, HDFS, AWS S3)
- Extensive experience in database management (NoSQL, MySQL, PostgreSQL)
- Proficiency in automation tools and frameworks (Ansible, Terraform, Kubernetes, Docker)
- Familiarity with modern CI/CD pipelines
- Programming skills in Python, Go, Java, or Scala
- Experience with monitoring and log management tools (Prometheus, Grafana, ELK Stack)
- Strong troubleshooting and debugging skills
- Excellent communication skills
- Bachelor’s degree or higher in Computer Science, Software Engineering, or related field
Tech stack
AWSGCPAzureKafkaHadoopSparkCassandraHDFSAWS S3NoSQLMySQLPostgreSQLAnsibleTerraformKubernetesDockerCI/CDPythonGoJavaScalaPrometheusGrafanaELK StackAerospikeSnowflake
Benefits
Commission eligibility for sales positionsBonus eligibility for non-sales positionsComprehensive benefits including physical, financial, and emotional supportPersonalized support options and expert guidance