Principal Software Engineer - Data Analytics at PubMatic | AdTechTalent
Engineering5 months agoHybrid
PubMatic
Principal Software Engineer - Data Analytics
JavaScalaPythonHadoopSparkKafkaSpark StreamingAWSSnowflakeREST APIBig DataGenAILLMLangChainPrompt EngineeringRAGDistributed SystemsAgileData AnalyticsBackend Development
Key details
Salary
Not specified
Employment type
Full-time
Seniority
Senior
Years experience
5-10
Location
Pune, India
Full job description
PubMatic is hiring a Principal Software Engineer focused on Data Analytics and AI agent development. The role involves building scalable big data platforms and pipelines using Hadoop, Spark, Kafka, Scala, and Snowflake, developing backend services with Java, REST APIs, JDBC, and AWS, and designing GenAI-powered analytics agents integrating LLMs like OpenAI and Claude. Responsibilities include leading projects, collaborating with cross-functional teams, managing GenAI workflows, and supporting customers. Requires 6+ years of Java/backend development experience, strong computer science fundamentals, expertise in Big Data tools, and proven GenAI application development skills. Bachelor’s degree in engineering or equivalent required. The position follows a hybrid work model with 3 days in office and 2 days remote. Benefits include parental leave, healthcare insurance, broadband reimbursement, and office amenities.
What you'll do
Build, design, and implement scalable, fault-tolerant big data platform processing terabytes of data
Develop backend services using Java, REST APIs, JDBC, and AWS
Similar jobs
More roles worth a look
Related opportunities based on specialty and working model so candidates can keep momentum.
Build and maintain Big Data pipelines with Spark, Hadoop, Kafka, and Snowflake
Architect and implement real-time data processing workflows and automation frameworks
Lead multiple projects developing features for data processing and reporting platforms
Collaborate with product managers and cross-functional teams to deliver end-to-end products and fix bugs
Design and develop GenAI-powered agents for analytics, operations, and data enrichment
Integrate LLMs (OpenAI, Claude, Mistral) into services for query understanding, summarization, and decision support
Manage end-to-end GenAI workflows including prompt engineering, fine-tuning, vector embeddings, and RAG
Improve availability and scalability of large data platforms and PubMatic software functionality
Participate in Agile/Scrum processes including sprint planning, retrospectives, backlog grooming, and prioritization
Communicate frequently with product managers about software features
Support customer issues via email or JIRA, provide updates and patches
Perform code and design reviews
Requirements
6+ years of coding experience in Java and backend development
Solid computer science fundamentals including data structure and algorithm design
Experience with full software development life cycle, coding standards, and code reviews
Hands-on experience with Big Data tools like Scala Spark, Kafka, Hadoop, and Snowflake
Proven experience in building GenAI applications including LLM integration, LangChain or similar, prompt engineering, embedding, and retrieval-based generation (RAG)
Experience developing and deploying scalable, production-grade AI or data systems
Ability to lead end-to-end feature development and debug distributed systems
Experience with large-scale big data pipelines, real-time systems, and data warehouses preferred
Ability to achieve stretch goals in a fast-paced environment
Ability to learn new technologies quickly and independently
Excellent verbal and written technical communication skills
Strong interpersonal skills and collaborative work ethic
Bachelor’s degree in engineering or equivalent from a recognized institute