DeepIntent seeks an MLOps Engineer for the European Data Infrastructure team. The role involves partnering with Data Science and AI teams to build and scale a machine learning and AI platform using Python, Spark, Kubernetes, Docker, and orchestration tools like Argo Workflows and Airflow. Responsibilities include adopting MLOps best practices, managing model tracking and versioning with MLflow, building CI/CD pipelines for ML artifacts, improving platform reliability and cost-efficiency, monitoring ML systems with Prometheus and Grafana, designing ML deployment infrastructure including GPU clusters, supporting LLM and generative AI workloads, collaborating on ML-driven feature development, and managing project priorities. Requirements include a Bachelor's degree or equivalent experience, strong software engineering skills (Python preferred), experience with Spark, Docker, Kubernetes, MLflow, Argo, Airflow, Kubeflow, Linux administration, familiarity with LLM/generative AI tooling and GPU infrastructure, and ability to work closely with data scientists. Benefits include competitive salary with bonus, medical insurance, flexible PTO, hybrid work options, professional development support, WiFi reimbursement, parental leave, and paid holidays.
What you'll do
Similar jobs
More roles worth a look
Related opportunities based on specialty and working model so candidates can keep momentum.
Continuously improve platform reliability, cost-efficiency and maintainability of the codebase
Establish monitoring and observability for ML/AI systems (Prometheus, Grafana) covering GPU utilization, model latency, throughput and custom ML metrics
Design and operate ML/AI deployment infrastructure including GPU cluster architecture, model serving and tool selection across training and inference workloads
Build and maintain infrastructure for LLM and generative AI workloads including model fine-tuning pipelines, vector databases, RAG architectures and inference optimization
Collaborate with business units and Product on ML/AI-driven feature development
Manage individual project priorities, deadlines and deliverables in a timely manner
Requirements
Bachelor's degree in Computer Science or similar technical field or equivalent practical experience
Strong software engineering skills in complex, distributed, multi-language systems (Python preferred)
Hands-on experience with Spark, Docker and Kubernetes in production environments
Experience building and operating end-to-end distributed systems
Experience developing and maintaining ML systems with open-source MLOps tools (MLflow, Argo, Metaflow, Airflow, Kubeflow)
Strong understanding of software testing, benchmarking, and CI/CD practices
Solid understanding of Linux systems administration
Familiarity with LLM/generative AI tooling and concepts (model serving frameworks, embeddings, vector stores, RAG pipelines) is a strong plus
Familiarity with GPU infrastructure - CUDA fundamentals, GPU scheduling/orchestration in Kubernetes, driver/toolkit management (NVIDIA drivers, CUDA toolkit, cuDNN)
Ability to work closely with data scientists and understand their tooling and workflows (Jupyter, notebooks, experiment tracking)
Enthusiastic learner with a genuine interest in ML/AI infrastructure
Competitive base salary plus performance-based bonusComprehensive medical insuranceFlexible PTOHybrid-friendly culture with flexible work optionsProfessional development reimbursementWiFi reimbursementParental leavePaid holidaysCareer development and advanced education supportWFH and internet stipends
Apply now
Ready to take the next step in your career? Click the button below to continue to the application process.