Senior Data Platform Engineer role to build and operate data infrastructure supporting products, analytics, machine learning, and AI systems. Responsibilities include managing Lakehouse architecture, distributed compute, data quality, governance, lineage, security, access controls, and platform operations. Collaborate with data engineers, scientists, and product teams. Requires 7+ years software engineering experience with Python, Scala, Java or JVM languages, Spark or distributed data-processing frameworks, production data platforms, and AWS services (EC2, S3, Athena, IAM). Experience with Kubernetes, Terraform, Airflow, Trino, ClickHouse, Starrocks, or Pinot is a plus. Benefits include unlimited PTO, medical/dental/vision coverage, equity, stock purchase plan, employee discounts, wellness classes, and pet insurance. Salary range $165,000 - $185,000. Hybrid role based in San Francisco, CA.
What you'll do
Build and operate Lakehouse, Spark, and distributed data-processing systems
Improve the performance, reliability, cost, and operational visibility of platform workloads
Similar jobs
More roles worth a look
Related opportunities based on specialty and working model so candidates can keep momentum.
Develop capabilities for data quality, governance, lineage, metadata, discoverability, security, and access control
Build the data context required for trusted reporting, ML models, and AI agents
Evaluate compute and query engines for read-heavy workloads, assessing cost, performance, reliability, and operational fit
Assess real-time OLAP technologies and identify appropriate production use cases
Improve deployment, observability, incident response, developer tooling, and platform operations
Partner with data scientists and data engineers to improve their experience using the platform
Contribute fixes and improvements to relevant open-source projects, including Spark, Iceberg, and Airflow
Mentor engineers through code reviews, pairing, design discussions, and knowledge-sharing sessions
Requirements
7+ years of software engineering experience
Strong experience in Python, Scala, Java, or another JVM language
Experience with Spark or another distributed data-processing framework
Experience building or operating production data platforms, backend systems, or distributed systems
Familiarity with at least one of these areas: Lakehouse architecture, Apache Iceberg, data quality, governance, lineage, metadata, query engines, or real-time OLAP
Working knowledge of AWS services such as EC2, S3, Athena, and IAM
Experience improving production reliability through monitoring, alerting, deployment automation, and incident response
Strong analytical, communication, and collaboration skills
Experience with Kubernetes, Terraform, Airflow, Trino, ClickHouse, Starrocks, or Pinot is a plus