Full job description
Senior ML Platform Engineer role at PubMatic to design and scale infrastructure and frameworks for machine learning development, experimentation, and production at petabyte scale. Responsibilities include building scalable ML pipelines, optimizing large-scale data workflows with Spark, Hadoop, Kafka, Snowflake, developing experiment tracking and observability frameworks, providing reusable AI/ML components, driving CI/CD and infrastructure-as-code for ML jobs, and collaborating cross-functionally. Requires 3-10 years experience with petabyte-scale datasets, proficiency in Python, Scala, Java, SQL, Spark, Hadoop, Kafka, Snowflake, ML Ops tools (Docker, Kubernetes, Airflow, MLflow), and strong understanding of programmatic advertising and Private Marketplace ecosystems. Bachelor’s or master's degree in computer science, data engineering or related field required.
What you'll do
- Design and maintain scalable ML pipelines and platforms for ingestion, feature engineering, training, evaluation, inference, and deployment
- Build and optimize large-scale data workflows using distributed systems (Spark, Hadoop, Kafka, Snowflake) to support analytics and model training
- Develop frameworks for experiment tracking, automated reporting, and observability to monitor model health, drift, and anomalies
- Work with industry-standard ML efficiency tools to optimize training workloads, accelerate experiments, and monitor performance at scale
- Provide reusable components, SDKs, and APIs that empower teams to leverage AI insights and ML models effectively
- Drive CI/CD, workflow orchestration, and infrastructure-as-code practices for ML jobs, ensuring reliability and reproducibility
- Partner cross-functional with Product, Data Science, and Engineering teams to align ML infrastructure with business needs
- Stay ahead of emerging trends in Generative AI, ML Ops, and Big Data to introduce best practices and next-gen solutions
- Work with petabyte-scale datasets and billions of transactions, powering global AdTech
- Apply AI/ML to deal troubleshooting, competitive intelligence, benchmarking, forecasting, and actionable insights
- Build advanced frameworks such as RAG systems, reinforcement learning strategies, and embedding platforms
- Convert business challenges into ML products, pioneering industry-first solutions
- Gain hands-on exposure to GPU/accelerated computing, Triton inference, and modern ML Ops frameworks
- Advance career with a clear growth path into applied ML engineering and research
- Be part of a culture that values experimentation, thought leadership, and cross functional collaboration
Requirements
- 3 to 10 years of hands on experience with petabyte-scale datasets and distributed systems
- Proficiency in SQL, Python (pandas, NumPy), and analytics/BI tools for data exploration and monitoring
- Strong understanding of Programmatic Advertising and Private Marketplace (PMP) ecosystems, including deal performance analytics, audience targeting, inventory optimization, and KPI driven marketplace optimization
- Familiarity with BI tools such as Looker or Grafana
- Passion for applied AI/ML and eagerness to bring research ideas into production
- Strong expertise with Spark, Hadoop, Kafka, and data warehousing (Snowflake, SparkSQL)
- Experience with CI/CD, Docker, Kubernetes, Airflow/MLflow, and experiment tracking tools
- Skilled in Python/Scala/Java for data-intensive applications
- Strong analytical skills and ability to debug complex data/ML pipeline issues
- Excellent communication and teamwork skills in cross-functional settings
- Bachelor’s or master's in computer science, Data Engineering, or related field
Tech stack
PythonScalaJavaSQLpandasNumPySparkHadoopKafkaSnowflakeSparkSQLLookerGrafanaDockerKubernetesAirflowMLflowTriton inferenceGPU/accelerated computing