Full job description
Design, develop, and manage advanced data structures and pipelines ensuring data quality and accessibility. Implement data solutions across platforms maintaining privacy and regulatory compliance. Act as a technical resource optimizing data processes and storage to support business initiatives. Requires strong SQL, Python, Apache Spark, cloud data warehouse management, and knowledge of AWS, GCP, or Azure. Familiarity with data pipeline design, data observability, and modern data stack tools like dbt, Airflow, Snowflake is highly desired. Responsibilities include designing data architectures, ensuring data quality, engineering ingestion frameworks, developing data consumption methods, implementing solutions on Kubernetes, Teradata, AWS, and Databricks, managing data lineage, collaborating cross-functionally, and working variable schedules including nights and weekends. Bachelor's degree or equivalent experience required with 5-7 years relevant experience.
What you'll do
- Design and construct data architectures and pipelines to standardize, transform, and ensure data integrity for business insights
- Guarantee data quality throughout ingestion, processing, and loading phases
- Engineer robust ingestion frameworks for various data types, monitor and report on data quality
- Develop accessible methods for data consumption including APIs, database views, and data extracts
- Implement data solutions on platforms such as Kubernetes, Teradata, AWS, and Databricks
- Select appropriate data storage platforms based on data sensitivity, privacy, and access needs
- Master data lineage and apply transformation rules for change management and issue resolution
- Collaborate with cross-functional teams to refine data sourcing and processing for optimal quality and efficiency
- Exercise independent judgment and discretion in significant matters
- Work nights and weekends as necessary with variable schedules
- Perform other duties as assigned
Requirements
- Strong SQL proficiency
- Programming language proficiency (Python preferred)
- Proficiency in distributed computing frameworks (e.g., Apache Spark)
- Cloud Data Warehouse management
- Knowledge of at least one cloud provider (AWS, GCP, or Azure)
- Familiarity with data pipeline design pattern
- Knowledge of data observability
- Exposure to Embedded Business Intelligence like Looker is desired but not required
- Exposure to modern data stack such as dbt, Airflow, Snowflake, Apache Spark is highly desired
- Exposure to Data Consumption challenges and solutions is highly desired
- Bachelor's Degree or equivalent combination of coursework and experience
- 5-7 years relevant work experience
Tech stack
SQLPythonApache SparkCloud Data WarehouseAWSGCPAzureLookerdbtAirflowSnowflakeKubernetesTeradataDatabricks
Benefits
Commission for most sales positionsBonus for most non-sales positionsBest-in-class benefits including physical, financial, and emotional supportPersonalized support options and expert guidance