As an Ontology Engineer on Samba TV's Knowledge Graph & Identity team, you will build, maintain, and extend knowledge graph schemas, derivation pipelines, and graph data models for measurement and audience intelligence products. Collaborate with Senior Ontologist and data scientists to implement ontological frameworks, contribute to entity resolution and data enrichment pipelines, and ensure graph accuracy and consistency. Write clean, production-quality Python and SPARQL code. Responsibilities include ontology implementation and validation, building SHACL validation shapes, supporting ontology versioning, developing event-to-ontology derivation pipelines using PySpark and Databricks, applying embedding-based approaches for entity matching, and collaborating with data engineering and product teams. Requirements include 2-4 years experience in knowledge graph or ontology engineering, proficiency with RDF, RDFS, OWL, SPARQL, SHACL, strong Python skills, data modeling knowledge, familiarity with entity resolution, and a relevant bachelor's degree. Preferred experience includes Amazon Neptune or Stardog, PySpark, embedding models, LLM APIs, and domain knowledge in media or ad tech.
What you'll do
Similar jobs
More roles worth a look
Related opportunities based on specialty and working model so candidates can keep momentum.
Implement and extend Samba's RDF/RDFS/OWL ontology schemas in the graph database adding entity classes, properties, and constraints under direction of Senior Ontologist
Build and maintain SHACL validation shapes for post-load graph consistency checks; identify and triage data quality and schema violations
Support ontology versioning, change log documentation, and consistency checking across schema updates
Write efficient, well-structured SPARQL queries and graph traversals to support downstream data science and product use cases
Contribute to event-to-ontology transformation and derivation layer by building PySpark/Databricks pipelines aggregating raw TV viewership and web activity events into durable graph attributes
Implement derivation logic specified by Senior Ontologist and data science team; validate outputs against SHACL shapes before graph load
Support incremental refresh and update logic aligned with graph's batch refresh cadence
Write production-quality Python code that is clean, well-tested, documented, and reusable
Work with PySpark and Databricks to process and transform high-volume data as part of graph pipeline development
Apply embedding-based approaches (semantic similarity, vector search) to entity matching and ontology alignment tasks
Contribute to team tooling, documentation, and reusable components that improve knowledge graph development efficiency
Partner closely with data engineering on pipeline design, data quality, and incremental ingestion patterns feeding the materialized graph substrate
Participate in ontology design reviews and cross-functional working groups
Work with product and operations teams to understand use case requirements and translate them into graph schema updates
Develop expertise in W3C semantic web standards, RDF-native graph databases, and entity resolution under guidance of Senior Ontologist
Requirements
2–4 years of hands-on experience in knowledge graph development, semantic data modeling, ontology engineering, or a closely related field
Working knowledge of W3C semantic web standards: RDF, RDFS, OWL, and SPARQL with practical experience querying or building in at least one triplestore or graph database
Familiarity with SHACL or equivalent constraint and validation frameworks for graph data quality
Strong Python skills - clean, readable, production-quality code with testing and documentation
Solid understanding of data modeling fundamentals - entity-relationship design, taxonomies, hierarchies, and how to represent complex real-world relationships in structured form
Familiarity with entity resolution or data matching concepts
Bachelor's degree required in Computer Science, Information Science, Mathematics, or a related field; Master's preferred
Detail-oriented and proactive about flagging data quality issues and schema inconsistencies
Hands-on experience with Amazon Neptune or Stardog or equivalent RDF-native triplestore; exposure to data virtualization a plus
Working knowledge of PySpark and Databricks for large-scale event aggregation and transformation pipelines
Familiarity with embedding models, vector search, or semantic similarity applied to entity matching, ontology alignment, or knowledge graph enrichment
Experience with LLM APIs or RAG-based approaches applied to information extraction, entity disambiguation, or schema mapping
Domain knowledge in media, entertainment, or ad tech - content metadata, advertising entities, TV viewership data, or audience/identity data
Exposure to identity resolution, probabilistic record linkage, or device graph approaches