AdTechTalent
Engineering21 days agoOn-site

Merkle

Lead ML Devops Engineer

GCPGoogle Cloud PlatformMLOpsVertex AIPythonCI/CDTerraformBigQueryCloud StoragePub/SubCloud MonitoringInfrastructure as CodeMachine LearningML DeploymentCloud EngineeringKubernetesKubeflowMLflow

Key details

Salary

Not specified

Employment type

Full-time

Seniority

Lead

Years experience

5-10

Location

Pune, India

Full job description

Seeking a Senior GCP MLOps Engineer with 5-8 years experience to automate deployment and lifecycle management of Python-based ML models on Google Cloud Platform. Responsibilities include designing scalable MLOps frameworks, automating deployment/testing/monitoring, establishing ML deployment standards, supporting model retraining and rollback, deploying models for batch and real-time inference, building CI/CD pipelines, implementing Infrastructure-as-Code with Terraform, managing cloud-native ML infrastructure including Vertex AI, BigQuery, Cloud Storage, Kubernetes Engine, Pub/Sub, and ensuring production reliability, scalability, and security. Requires strong GCP hands-on experience, Python ML deployment expertise, CI/CD pipeline development, monitoring and troubleshooting production ML workloads. Preferred certifications include Google Cloud Professional ML Engineer and familiarity with MLflow or Kubeflow.

What you'll do

  • Design, build, and maintain scalable MLOps frameworks on Google Cloud Platform
  • Automate deployment, testing, monitoring, and lifecycle management of machine learning models
  • Establish repeatable and standardized ML deployment processes across environments
  • Implement model versioning, artifact management, and deployment governance standards
  • Support model retraining, rollback, and release management processes
  • Deploy Python-based machine learning models into production environments
  • Build automated deployment pipelines for batch and real-time inference workloads
  • Develop reusable deployment templates and automation frameworks
  • Support model serving using Vertex AI Endpoints and containerized deployment architectures
  • Ensure high availability, reliability, and scalability of production ML services
  • Design and implement CI/CD pipelines for machine learning applications and services
  • Integrate source control, testing, and deployment workflows into enterprise delivery pipelines
  • Implement Infrastructure-as-Code (IaC) practices for repeatable environment provisioning
  • Support environment management across development, testing, and production environments
  • Design and support cloud-native ML infrastructure on GCP
  • Manage and optimize services including Vertex AI, Cloud Storage, BigQuery, Cloud Build, Cloud Run, Kubernetes Engine (GKE), Pub/Sub
  • Optimize infrastructure for performance, reliability, security, and cost efficiency
  • Troubleshoot production issues and support platform stability initiatives
  • Implement monitoring and alerting frameworks for deployed machine learning services
  • Track model performance, operational health, latency, and system utilization
  • Support model lifecycle governance and operational compliance requirements
  • Establish logging, observability, and operational dashboards
  • Drive best practices for production support and operational excellence

Requirements

  • Bachelor's degree in Computer Science, Engineering, Information Technology, or related discipline
  • 5 - 8 years of experience in Cloud Engineering, MLOps, or ML Platform Engineering
  • Strong hands-on experience with Google Cloud Platform (GCP)
  • Proven experience deploying and operationalizing Python-based machine learning models
  • Strong experience with Vertex AI and production ML deployment patterns
  • Experience building CI/CD pipelines for machine learning applications
  • Experience implementing Infrastructure-as-Code using Terraform or similar tools
  • Experience monitoring and supporting production machine learning workloads
  • Strong troubleshooting and problem-solving skills
  • Preferred: Google Cloud Professional Machine Learning Engineer Certification
  • Preferred: Familiarity with MLflow, Kubeflow, or similar MLOps frameworks

Tech stack

Google Cloud PlatformGCPVertex AIPythonCloud BuildGitHub ActionsJenkinsGitLab CI/CDTerraformBigQueryCloud StoragePub/SubCloud MonitoringLoggingAlertingGitGitHubMLflowKubeflow

Apply now

Ready to take the next step in your career? Click the button below to continue to the application process.

Similar jobs

More roles worth a look

Related opportunities based on specialty and working model so candidates can keep momentum.