Full job description
Seeking a Software Engineer with 1-3 years experience in scalable, distributed environments to join Epsilon's Data Practice team. Responsibilities include designing, developing, and maintaining scalable data processing solutions using Python, Apache Spark, and Databricks on-premises and cloud (AWS). Optimize Spark jobs, build fault-tolerant data pipelines with monitoring and alerting, develop SQL queries, and collaborate with cross-functional teams. Required skills: Scala, Python, or Java programming; Apache Spark; Hadoop and Hive exposure; cloud platform experience (AWS preferred); Kubernetes, Docker, and Airflow exposure; strong problem-solving and communication skills; BA/BS in Computer Science or related field. Machine Learning experience and Databricks certifications are a plus. Location: Westminster, Colorado, USA. Salary range: $73,500 - $136,500 annually. Benefits include flexible time off, paid holidays, sick time, parental leave, health coverage, 401(k), tuition assistance, and more.
What you'll do
- Design, develop, and maintain scalable data processing solutions on-premises and cloud
- Optimize and fine-tune Spark jobs for performance and resource utilization
- Design and implement scalable, fault-tolerant data pipelines with monitoring, alerting, and logging
- Leverage AWS and Databricks cloud services to build and manage data pipelines and distributed workloads
- Develop and optimize SQL queries across relational and data warehouse systems
- Apply design patterns and best practices for data modeling, partitioning, and distributed system performance
- Use Git or equivalent for source control and maintain unit and integration testing
- Collaborate with Product Owners and cross-functional teams to translate business requirements into technical solutions
- Extract actionable insights from large datasets to support data-driven decision-making
Requirements
- 1-3 years of software development experience in scalable, distributed, or multi-node environments
- Proficient in Scala, Python, or Java
- Experience with Apache Spark and exposure to Hadoop, Hive, and related big data technologies
- Experience with cloud platforms, preferably AWS
- Exposure to Kubernetes, Docker, and MWAA/Airflow
- Strong problem-solving skills and ability to own problems end-to-end
- Consultative mindset with strong communication and collaboration skills
- BA/BS in Computer Science or related discipline
- Experience in Machine Learning including model development, feature engineering, or integrating ML workflows
- Experience with Databricks tools such as Notebooks, Delta Lake, Jobs, Pipelines, and Unity Catalog preferred
- AWS, Databricks, or Spark certification is a plus
Tech stack
PythonApache SparkDatabricksScalaJavaHadoopHiveAWSKubernetesDockerMWAAAirflowSQLDelta LakeGit
Benefits
Flexible time off (FTO)15 paid holidaysPaid sick timeParental/new child leaveChildcare & elder care assistanceAdoption assistanceComprehensive health coverage401(k) planTuition assistanceCommuter benefitsProfessional developmentEmployee recognitionCharitable donation matchingHealth coaching and counseling