Full job description
Microsoft Advertising seeks a Senior Principal Architect to lead the design and development of the Ads Trust & Safety AI Platform. This role involves setting technical direction, shaping platform strategy, and building scalable systems for Trust & Safety, Risk, Fraud, Security, Policy, and Enforcement. Responsibilities include defining platform architecture, translating business needs into reusable capabilities, establishing design principles and standards, and collaborating with cross-functional teams. The candidate will architect AI-powered investigation workflows, entity intelligence systems, decisioning and enforcement platforms, and adversarial behavior detection systems. The role requires 15+ years of software engineering experience, 8+ years in senior technical leadership, and deep expertise in AI/ML systems, model serving, distributed systems, and Trust & Safety domains. Strong communication and ability to influence multiple teams are essential. Preferred qualifications include experience with adversarial systems, deep research agents, heterogeneous inference platforms, knowledge graphs, human-in-the-loop review systems, and collaboration with industry partners.
What you'll do
- Define long-term architecture for Ads Trust & Safety AI Platform covering ingestion, signal acquisition, entity intelligence, retrieval, model orchestration, agentic workflows, decisioning, enforcement, human review, audit, and measurement
- Translate Trust & Safety, Risk, Fraud, Security, and Policy needs into reusable platform capabilities
- Establish reference architectures, design principles, technical standards, and engineering patterns for high-integrity AI and decisioning systems
- Drive architecture choices balancing latency, throughput, quality, cost, explainability, governance, reliability, and operational safety
- Identify platform gaps and create roadmap balancing near-term delivery with long-term leverage
- Architect platform for Deep Research Agents investigating domains, landing pages, advertisers, business entities, ownership patterns, web presence, reputation, policy risk, and fraud signals
- Architect workflows combining retrieval, crawling, structured evidence extraction, LLM reasoning, policy grounding, risk scoring, and human-in-the-loop
- Architect guardrails for agentic systems including source provenance, confidence scoring, hallucination controls, audit logs, escalation paths, and human override
- Partner with Applied Science to convert AI research prototypes into production systems meeting quality, latency, cost, reliability, safety, and governance targets
- Architect systems for high-fidelity understanding of domains, websites, landing pages, advertisers, business identities, ownership structures, relationship graphs, reputation, and provenance
- Design real-time, nearline, and batch scoring systems for policy enforcement, fraud detection, abuse prevention, advertiser risk scoring, and marketplace protection
- Evolve abstractions for model orchestration, feature lookup, signal stores, retrieval, model versioning, decision logging, policy controls, fallbacks, and experimentation
- Architect systems to detect, learn, and mitigate adversarial behavior across advertiser lifecycle including account creation, login events, payment changes, budget changes, campaign edits, creative changes, landing-page changes, and enforcement history
- Build sequential and event-based risk systems reasoning over advertiser behavior over time
- Collaborate across Microsoft teams in Safety, Security, Responsible AI, Identity, and partners to create shared platform capabilities for risk detection, abuse prevention, evidence generation, and enforcement governance
- Evangelize Trust & Safety AI platform strategy across Microsoft to converge on shared architectures, reusable abstractions, common taxonomies, and consistent decisioning patterns
- Evangelize and learn from industry peers and approved partner networks facing similar abuse patterns and translate learnings into platform improvements
Requirements
- Bachelor’s Degree in Computer Science or related technical field or equivalent practical experience
- 15+ years of professional software engineering experience
- 8+ years of senior technical leadership experience
- Proven experience architecting and delivering large-scale production systems with reliability, scalability, latency, correctness, availability, security, and operational requirements
- Deep technical experience in AI/ML systems, agentic systems, model-serving infrastructure, decisioning systems, distributed systems, data platforms, workflow platforms, risk platforms, security platforms, or Trust & Safety systems
- Experience building or leading production AI/ML systems including LLM-based workflows, retrieval-augmented generation, model orchestration, automated reasoning, human-in-the-loop systems, AI-assisted operational tooling, or agentic workflows
- Strong understanding of engineering requirements for deploying AI or decision systems in production including evaluation, observability, quality measurement, rollout safety, fallback behavior, latency/cost tradeoffs, drift detection, explainability, governance, and operational reliability
- Experience designing high-integrity systems with auditable, reproducible, explainable, governed, and secure decisions
- Ability to drive clarity from ambiguity, define technical direction, create reusable platform abstractions, and influence execution across multiple teams without direct management authority
- Strong written and verbal communication skills to explain architecture, tradeoffs, risks, sequencing, and technical strategy to senior leaders
- Preferred: experience in Trust & Safety, Fraud, Abuse, Risk, Security, Ads Quality, Marketplace Integrity, Policy Enforcement, or advertiser protection systems
- Preferred: experience with adversarial systems such as phishing, malware, cloaking, account takeover, payment abuse, fake identities, coordinated fraud, and policy evasion
- Preferred: experience building Deep Research Agents, investigation agents, reviewer-assist systems, retrieval-augmented generation systems, LLM-powered operational workflows, or AI systems producing grounded evidence
- Preferred: expertise in heterogeneous inference platforms supporting LLMs, SLMs, wide & deep models, ensembles, graph models, classical ML models, heuristics, and rules engines
- Preferred: experience with entity intelligence, knowledge graphs, web crawling, domain reputation, business identity resolution, provenance, evidence extraction, or risk scoring
- Preferred: experience designing human-in-the-loop review systems, appeals workflows, audit platforms, policy reasoning systems, or enforcement governance mechanisms
- Preferred: experience with large-scale measurement systems for false positives, false negatives, model drift, agent quality, policy quality, reviewer quality, enforcement stability, business impact, and operational health
- Preferred: experience collaborating with Trust, Safety, Security, Privacy, Identity, Compliance, Legal, or Responsible AI teams across multiple products or platforms
- Preferred: experience evangelizing technical strategy across multiple teams, learning from industry peers, and helping establish shared standards, taxonomies, schemas, signal-quality measures, or platform patterns
- Preferred: experience working with industry partners, trusted abuse-prevention networks, threat-intelligence providers, domain-reputation providers, identity-verification providers, payment-risk partners, or ecosystem safety initiatives
Tech stack
AIMLLLMmodel-serving infrastructuredistributed systemsdata platformsworkflow platformsrisk platformssecurity platformsTrust & Safety systemsretrieval-augmented generationmodel orchestrationautomated reasoninghuman-in-the-loop systemsagentic workflowsknowledge graphsweb crawlingpolicy enforcement systemsaudit platformspolicy reasoning systemsheterogeneous inference platformsgraph modelsclassical ML modelsheuristicsrules engines