AdTechTalent
Data Science18 days agoOn-site

Samba TV

Data Scientist

pythonpysparkdatabricksdelta lakesqlawsgcpairflowmlopsmachine learningairag systemsllmvector databasessemantic searchdata sciencedataopsentity resolutionprobabilistic record linkageembedding-based matchingcausal inferencea/b testingsynthetic controluplift modelingmediaad techmeasurementaudience modeling

Key details

Salary

Not specified

Employment type

Full-time

Seniority

Mid-level

Years experience

3-5

Location

Warsaw, Poland

Full job description

Mid-level Data Scientist role in Warsaw responsible for end-to-end delivery of data science projects with minimal guidance. Requires deep expertise in measurement or audience modelling and ability to build production-ready ML and AI solutions. Responsibilities include project ownership, methodology decisions, solution design, coding in Python and PySpark on Databricks, developing reusable tools, mentoring juniors, and cross-functional collaboration. Requires Bachelor's degree (Master's preferred) in quantitative field, 3-5 years experience, advanced Python, SQL, PySpark skills, knowledge of Databricks, Delta Lake, cloud platforms (AWS/GCP), core ML techniques, MLOps practices, and exposure to modern AI methods. Preferred skills include knowledge graph construction, probabilistic record linkage, causal inference, and media/ad tech experience.

What you'll do

  • Own end-to-end delivery of significant data science projects from problem scoping to production deployment
  • Make independently-reasoned decisions on methodology, model selection, and evaluation; document technical solutions
  • Lead solution design; break down complex epics into user stories with clear acceptance criteria
  • Adopt DataOps and MLOps best practices including experiment tracking, pipeline orchestration, model monitoring, reproducibility
  • Build production-quality Python and PySpark code on Databricks; implement advanced ML and AI workflows
  • Develop and maintain reusable tools, libraries, and documentation to improve team efficiency and standards
  • Conduct code reviews with constructive feedback
  • Mentor junior data scientists on technical execution, code quality, and career development
  • Lead internal talks or workshops on ML topics
  • Collaborate cross-functionally with product, engineering, and operations; translate business requirements into technical specifications
  • Partner with data engineering on scalable pipeline design
  • Participate in cross-functional design reviews and working groups

Requirements

  • Bachelor's degree in Statistics, Data Science, Computer Science, Mathematics or related quantitative field; Master's preferred
  • 3–5 years of hands-on data science experience with ability to own and deliver complex projects independently
  • Advanced Python with production-quality code, testing, and documentation
  • Strong SQL and PySpark for billion-row datasets
  • Experience with Databricks workflows, Delta Lake, and job orchestration
  • Working knowledge of cloud platforms (AWS or GCP)
  • Solid command of core ML techniques: regression, classification, clustering, model evaluation, experimental design
  • Proficiency with MLOps practices: experiment tracking, pipeline orchestration (Airflow), reproducible model deployment
  • Exposure to modern AI methodologies: RAG systems, LLM-augmented models, vector databases, semantic search
  • Strong communication skills for documentation and cross-functional collaboration
  • Demonstrated ability to mentor junior data scientists

Tech stack

PythonPySparkDatabricksDelta LakeSQLAWSGCPAirflowMLAIRAG systemsLLM-augmented modelsvector databasessemantic searchRDFOWLSPARQL

Apply now

Ready to take the next step in your career? Click the button below to continue to the application process.

Similar jobs

More roles worth a look

Related opportunities based on specialty and working model so candidates can keep momentum.