Digital Turbine seeks a Principal Software Engineer to lead enterprise-wide ML and AI architecture. The role involves defining a 2-3 year vision for ML platforms, solving complex technical challenges in model scaling and distributed systems, evaluating ML technologies, and establishing MLOps standards. Responsibilities include mentoring senior engineers, leading architecture governance, fostering engineering culture, architecting high-scale ML infrastructure, integrating generative AI and LLM technologies, designing observability systems, and collaborating with executive and cross-functional teams. Requirements include a Bachelor’s degree in a quantitative field (Master’s or Ph.D. preferred), 10+ years of software/ML engineering experience (or 8+ with advanced degree), expertise in PyTorch, TensorFlow, Jax, Kubernetes, Python, C++, Rust, and experience with cloud-native infrastructure and MLOps tools. Preferred qualifications include expertise in generative AI, distributed computing, hardware acceleration, and industry leadership. The position is hybrid and limited to candidates local to New York, NY or Berlin, Germany.
What you'll do
Define enterprise architecture and lead long-term vision and technology roadmap for high-scale ML platforms and generative AI systems
Similar jobs
More roles worth a look
Related opportunities based on specialty and working model so candidates can keep momentum.
Solve novel, ambiguous engineering challenges in model scaling, distributed systems, real-time inference, and feature engineering
Evaluate frontier ML research, frameworks, hardware accelerators, and third-party platforms for strategic technology choices
Establish company-wide MLOps best practices, security standards, evaluation frameworks, model governance protocols, and operational reliability metrics
Mentor and sponsor Staff, Senior, and peer engineers to elevate technical bar, design rigor, and execution speed
Lead Architecture Review Boards approving critical system designs, data pipelines, and production deployment architectures
Foster a culture of technical excellence, continuous learning, operational resilience, and principled engineering tradeoffs
Architect and optimize large-scale distributed training clusters, feature platforms, high-throughput model serving engines, and cost-efficient inference pipelines
Lead design and integration of advanced foundation models, LLM orchestration, retrieval-augmented generation, parameter-efficient fine-tuning, and vector infrastructure
Design enterprise-grade observability systems to monitor model drift, system health, data quality, security posture, and business metric impact
Partner with VPs, Directors, Product Leaders, and Domain Experts to translate business vision into scalable technical roadmaps and architectural specs
Communicate complex technical concepts, trade-offs, risks, and strategic investments to executive leadership and non-technical stakeholders
Build operational bridges between ML Engineering, Data Engineering, Infrastructure/DevOps, Enterprise Security, and Product teams
Requirements
Bachelor’s degree in Computer Science, Machine Learning, Data Science, or related quantitative field (Master’s or Ph.D. preferred)
10+ years of software/ML engineering experience with a Bachelor’s degree, OR 8+ years with an advanced degree
Proven track record operating as a Principal (P5) or Staff (P4) level engineer on enterprise-scale systems
Expert-level mastery of PyTorch, TensorFlow, Jax, and modern distributed ML training frameworks
Comprehensive mastery of cloud-native infrastructure, container orchestration (Kubernetes), feature stores, CI/CD pipelines, and enterprise MLOps suites
Exceptional proficiency in system architecture, distributed computing, memory optimization, data structures, and languages such as Python, C++, or Rust
Demonstrated ability to drive strategic alignment, architectural consensus, and engineering compliance across non-reporting teams
Ability to balance immediate execution needs with long-term architectural stability, scalability, and cost efficiency