AdTechTalent
Data Science1 month agoHybrid

Epsilon

Senior Software Engineer (Big Data)

pythonapache sparkscalajavahadoophiveawsdatabrickssqlmachine learningelk stackkubernetesdockerairflowdata pipelinesbig datadistributed systemsdata science

Key details

Salary

$89K – $165K

Employment type

Full-time

Seniority

Senior

Years experience

5-10

Location

Westminster, United States

Full job description

Design, develop, and maintain scalable data processing solutions using Python and Apache Spark across on-premises and cloud environments. Optimize Spark jobs for performance and resource utilization. Build scalable, fault-tolerant data pipelines with monitoring, alerting, and logging. Use AWS cloud services to manage data pipelines and distributed workloads. Develop and optimize SQL queries for relational and data warehouse systems. Apply best practices for data modeling and distributed system performance. Use Git for source control and maintain unit and integration testing. Collaborate with teams to translate business requirements into technical solutions. Extract insights from large datasets to support decision-making. Mentor junior engineers and conduct code reviews. Requirements include 5+ years experience in scalable distributed environments, proficiency in Scala, Python, or Java, experience with Apache Spark, Hadoop, Hive, AWS, Databricks, Kubernetes, Docker, Airflow, and machine learning. BA/BS in Computer Science or related field required. Salary range: USD $88,900 - $165,100.

What you'll do

  • Design, develop, and maintain scalable data processing solutions across on-premises and cloud environments
  • Optimize and fine-tune Spark jobs for performance including resource utilization, shuffling, partitioning, and caching
  • Design and implement scalable, fault-tolerant data pipelines with monitoring, alerting, and logging
  • Leverage AWS cloud services to build and manage data pipelines and distributed processing workloads
  • Develop and optimize SQL queries across relational and data warehouse systems
  • Apply design patterns and best practices for data modeling, partitioning, and distributed system performance
  • Use Git for source control and maintain strong unit and integration testing practices
  • Collaborate with Product Owners and multi-functional teams to translate business requirements into technical solutions
  • Extract actionable insights from large datasets to support data-driven decision-making
  • Mentor junior engineers, conduct code reviews, and contribute to engineering best practices and standards

Requirements

  • 5+ years of software development experience in scalable, distributed, or multi-node environments
  • Proficient in Scala, Python, or Java
  • Significant experience with Apache Spark
  • Exposure to Hadoop, Hive, and related big data technologies
  • Experience with AWS cloud platform preferred
  • Experience with Databricks platform preferred
  • Experience with modern data tools such as Kubernetes, Docker, and Airflow
  • Strong problem-solving skills
  • Consultative mentality and strong communication skills
  • BA/BS in Computer Science or related field
  • Experience in Machine Learning including model development or integrating ML workflows
  • Experience with Databricks Notebooks, Delta Lake, Jobs, Pipelines, Unity Catalog preferred
  • Proficiency with ELK stack preferred
  • AWS, Databricks, or Spark certification a plus

Tech stack

PythonApache SparkScalaJavaHadoopHiveAWSDatabricksSQLRDBMSData WarehouseGitKubernetesDockerAirflowELK stackElasticsearchLogstashKibanaDelta LakeUnity Catalog

Benefits

Flexible time off (FTO)15 paid holidaysPaid sick timeParental/new child leaveChildcare & elder care assistanceAdoption assistanceComprehensive health coverage401(k)Tuition assistanceCommuter benefitsProfessional developmentEmployee recognitionCharitable donation matchingHealth coaching and counseling

Apply now

Ready to take the next step in your career? Click the button below to continue to the application process.

Similar jobs

More roles worth a look

Related opportunities based on specialty and working model so candidates can keep momentum.