Digital Turbine seeks a Principal Software Engineer to lead enterprise-wide ML and AI architecture. Responsibilities include defining long-term ML platform vision, solving complex technical challenges in model scaling and real-time inference, evaluating ML technologies, establishing MLOps standards, mentoring senior engineers, leading architecture governance, and fostering engineering culture. The role involves architecting large-scale distributed training and inference systems, integrating generative AI and LLM technologies, designing observability systems, and collaborating with executive leadership and cross-functional teams. Requirements include a Bachelor’s degree in a quantitative field (Master’s or Ph.D. preferred), 10+ years of software/ML engineering experience (or 8+ with advanced degree), expertise in PyTorch, TensorFlow, Jax, Kubernetes, cloud-native infrastructure, Python/C++/Rust programming, and leadership skills in strategic alignment and architectural governance. Location is hybrid in New York, USA or Berlin, Germany. Candidates must be local to the posting location.
What you'll do
Define the 2–3 year vision for ML architecture and build foundational platform capabilities
Similar jobs
More roles worth a look
Related opportunities based on specialty and working model so candidates can keep momentum.
Solve complex, novel technical problems in model scaling, distributed systems, real-time inference, and feature engineering
Evaluate frontier ML research, frameworks, hardware accelerators, and third-party vendor platforms for strategic technology choices
Establish company-wide MLOps best practices, security standards, evaluation frameworks, model governance protocols, and operational reliability metrics
Mentor senior and staff engineers, elevating technical bar, design rigor, and execution speed
Lead Architecture Review Boards (ARBs) to approve critical system designs and production deployment architectures
Foster a culture of technical excellence, continuous learning, operational resilience, and principled engineering tradeoffs
Architect and optimize large-scale distributed training clusters, feature platforms, high-throughput model serving engines, and cost-efficient inference pipelines
Lead design and integration of advanced foundation models, LLM orchestration, retrieval-augmented generation, parameter-efficient fine-tuning, and vector infrastructure
Design enterprise-grade observability systems to monitor model drift, system health, data quality, security posture, and business metric impact
Partner with VPs, Directors, Product Leaders, and Domain Experts to translate business vision into scalable technical roadmaps and architectural specs
Communicate complex technical concepts, trade-offs, risks, and strategic investments to executive leadership and non-technical stakeholders
Build operational bridges between ML Engineering, Data Engineering, Infrastructure/DevOps, Enterprise Security, and Product teams
Requirements
Bachelor’s degree in Computer Science, Machine Learning, Data Science, or related quantitative field (Master’s or Ph.D. preferred)
10+ years of software/ML engineering experience with a Bachelor’s degree, OR 8+ years with an advanced degree
Proven track record operating as a Principal (P5) or Staff (P4) level engineer on enterprise-scale systems
Expert-level mastery of PyTorch, TensorFlow, Jax, and modern distributed ML training frameworks
Comprehensive mastery of cloud-native infrastructure, container orchestration (Kubernetes), feature stores, CI/CD pipelines, and enterprise MLOps suites
Exceptional proficiency in system architecture, distributed computing, memory optimization, data structures, and languages such as Python, C++, or Rust
Demonstrated ability to drive strategic alignment, architectural consensus, and engineering compliance across non-reporting teams and business units
Ability to balance immediate execution needs with long-term architectural stability, scalability, and cost efficiency