data engineeringapache sparkpysparksqlazuredatabricksdata lakepower bitableauapi ingestiondata governancerbacabacworkflow orchestrationapache airflowcontainerizationkubernetesc4 modelplantuml
Key details
Salary
Not specified
Employment type
Full-time
Seniority
Mid-level
Years experience
3-5
Location
Bengaluru, India
Full job description
Looking for a Data Engineer to build and maintain data platforms ingesting, transforming, and serving data for media and sustainability analytics products. Responsibilities include building data pipelines from APIs and internal sources into cloud data lakes, developing Spark/PySpark and SQL transformations, maintaining orchestration workflows, supporting data governance and access control, implementing platform features, integrating with BI tools, contributing to architecture documentation, troubleshooting pipeline issues, collaborating with DevOps/security teams, and participating in on-call rotations. Requires 3+ years experience with Apache Spark, SQL, Azure cloud, Databricks, Unity Catalog, data governance models (RBAC/ABAC), BI tool integration, API ingestion tools, and software engineering best practices. Nice to have experience with Apache Airflow, identity/access management (Okta, Entra ID), service mesh (Istio), containerized deployments (AKS/Kubernetes), sustainability data models, and C4 model documentation.
What you'll do
Build and maintain data pipelines that ingest data from third-party APIs and internal sources into cloud data lake and lakehouse environments
Similar jobs
More roles worth a look
Related opportunities based on specialty and working model so candidates can keep momentum.
Develop data transformation logic (Spark/PySpark, SQL) to standardize, model, and enrich raw data into analytics-ready datasets
Build and maintain orchestration workflows to schedule, monitor, and troubleshoot data pipeline execution
Support data governance and access control models (e.g. Unity Catalog, ABAC-based policies) to help ensure data is secure and appropriately scoped by tenant, client, or market
Work with product managers and senior engineers to implement platform features such as connector frameworks, taxonomy/rules engines, and data export capabilities
Support integration with visualization and reporting tools (e.g. Power BI, Tableau) and help ensure downstream data consumers have reliable, well-documented access
Contribute to architecture documentation (e.g. C4 model diagrams) and participate in design reviews
Troubleshoot data quality, pipeline failures, and performance issues, tracing errors from source to destination
Work with DevOps/security teams on service account management, credential handling, and infrastructure migrations (e.g. containerization)
Participate in on-call/support rotations as needed for production data pipelines
Requirements
3+ years of experience as a Data Engineer building production-grade data pipelines
Solid hands-on experience with Apache Spark (PySpark) and SQL for data transformation at scale
Experience with cloud platforms (Azure preferred) and cloud-native data storage (e.g. Data Lake / Blob Storage)
Experience with Databricks, including familiarity with Unity Catalog or similar data governance/catalog tools
Familiarity with data governance and access control models (RBAC/ABAC), and working with sensitive, multi-tenant data
Experience integrating data pipelines with BI/visualization tools (Power BI, Tableau, or similar)
Comfortable working with API-based data ingestion tools/connectors (e.g. Adverity or similar ingestion platforms) is a plus
Solid understanding of software engineering practices: version control, CI/CD, testing, code review
Good communication skills and ability to work cross-functionally with product, engineering, and client-facing stakeholders