Full job description
PubMatic seeks a senior engineer specializing in Generative AI and AI agent development. Responsibilities include leading design, development, and deployment of AI-driven features using Retrieval-Augmented Generation (RAG), vector databases, and large language models (LLMs). The role involves optimizing LLMs, developing AI agents, designing vector databases (FAISS, Pinecone, Weaviate), prompt engineering, and performance evaluation. Candidates should have 2-10 years of experience with LLMs, AI agents, agentic frameworks (LangGraph, CrewAI, AutoGen), vector databases, observability tools (Langfuse), Python, TensorFlow, PyTorch, and Hugging Face Transformers. A bachelor's degree in engineering or equivalent is required. The position is full-time with a hybrid work schedule based in Pune, India. Benefits include parental leave, healthcare insurance, broadband reimbursement, and office amenities.
What you'll do
- Lead design, development, and deployment of AI-driven features with end-to-end ownership
- Spearhead technical design meetings and produce detailed design documents for scalable, secure AI architectures
- Align solutions with long-term product strategy and technical roadmaps
- Implement and optimize LLMs including fine-tuning, deploying pre-trained models, and evaluating performance
- Develop AI agents powered by RAG systems integrating external knowledge sources
- Design, implement, and optimize vector databases for efficient and scalable vector search
- Create and fine-tune sophisticated prompts to improve LLM performance
- Utilize evaluation frameworks and metrics to assess and improve generative models and AI systems
- Collaborate with data scientists, engineers, and product teams to integrate AI capabilities into products and tools
- Stay updated with latest research and trends in LLMs, RAG, and generative AI
- Continuously monitor and optimize models for performance, scalability, and cost efficiency
Requirements
- 2 to 10 years of experience and strong understanding of LLMs and their underlying principles — transformer architecture, attention mechanisms, and hyperparameter tuning
- Proven experience designing and building AI agents, including multi-agent orchestration, tool-use patterns, multi-step planning, and agent memory architectures
- Hands-on experience with agentic frameworks such as LangGraph, CrewAI, or AutoGen, and familiarity with RAG pipelines integrating external knowledge sources
- In-depth knowledge of vector databases and indexing algorithms; practical experience with FAISS, Pinecone, Weaviate, or Milvus
- Experience with agent observability, tracing, and guardrails using tools like Langfuse or equivalent
- Proficiency in prompt engineering for context-sensitive, domain-specific LLM outputs
- Familiarity with Evals and other performance evaluation tools
- Proficiency in Python and experience with machine learning libraries such as TensorFlow, PyTorch, and Hugging Face Transformers
- Experience with data preprocessing, vectorization, and handling large-scale datasets
- Ability to present complex technical ideas and results to both technical and non-technical stakeholders
- Bachelor’s degree in engineering or equivalent from a well-known institute/university
Tech stack
Generative AIAI agentsRetrieval-Augmented Generation (RAG)vector databaseslarge language models (LLMs)FAISSPineconeWeaviateMilvusLangGraphCrewAIAutoGenLangfusePythonTensorFlowPyTorchHugging Face TransformersEvalsDockerKubernetesAWSGCPAzure
Benefits
Paternity/maternity leaveHealthcare insuranceBroadband reimbursementKitchen with healthy snacks and drinksCatered lunches