Senior Principal AI Architect/Engineer
About the role
The AI Observability Architect is a senior technical leader responsible for designing, deploying, and operating an enterprise-grade, production-ready AI observability platform that spans the full spectrum of modern agentic AI—from large language model (LLM) workflows and multi-agent orchestration to physical AI systems, reinforcement learning harnesses, multi-modal pipelines, and agentic marketplaces. This role serves as the strategic and engineering authority for end-to-end telemetry, tracing, safety, and quality signals across heterogeneous agent frameworks and platforms. The architect leads the convergence of AI observability with safety & security (including red teaming), Responsible AI (RAI), data science, physical AI, memory/skills engineering, agent fleet management, self-evolving harnesses, reinforcement learning, agent-to-agent protocols (A2A, UCP, AP2), and continuous quality engineering—making this a uniquely broad and high-impact role within the AI Solutions & Platforms organization. The role also owns OpenTelemetry (OTEL) integration across third-party agentic platforms (Salesforce AgentForce, ServiceNow, Microsoft Agent 365, and others), enabling unified observability and governance at enterprise scale.
Responsibilities
- Agentic AI Observability Architecture at Scale (30%)
- Define and own the enterprise observability architecture for AI agents, LLMs, multi-agent workflows, and physical AI systems—covering planner/executor loops, tool/function calls, RAG retrieval chains, and memory/state transitions.
- Build and operate unified telemetry pipelines incorporating metrics, logs, distributed traces, semantic/vector signals, and real-time event streaming (Kafka) at enterprise scale.
- Instrument OpenTelemetry (OTEL) across heterogeneous platforms including Salesforce AgentForce, ServiceNow, Microsoft Agent 365, and internal frameworks—delivering protocol-level observability for agent ecosystems including MCP, A2A, UCP, and AP2.
- Design and implement observability for Agent Fleets, multi-modal pipelines, physical AI systems, and self-evolving reinforcement learning harnesses—including signal capture for reward shaping and policy evaluation.
- Deliver dashboards, alerting, SLO/SLA management, incident runbook automation, and RCA tooling that drive measurable reliability improvements and reduce MTTR across agentic services.
- Establish cost telemetry and FinOps observability for AI workloads—token consumption, inference cost allocation, and GPU/compute efficiency across cloud environments (Azure, AWS, GCP).
- Safety, Security & Red Teaming (15%)
- Lead observability-driven red team exercises targeting agentic AI systems—instrumenting attack surfaces, adversarial prompt injection vectors, model evasion attempts, and multi-agent trust boundary failures.
- Design telemetry pipelines that capture safety-critical signals: guardrail trigger rates, policy violation events, PII exposure risks, prompt leakage, and agent hallucination rates.
- Partner with Security and RAI teams to embed threat modeling, zero-trust agent authentication, and behavioral anomaly detection into the observability platform.
- Instrument secure policy enforcement layers across agent-to-agent communication protocols (A2A, UCP, AP2) and maintain audit-ready traceability for all AI decision events.
- Develop and maintain a Security Observability Playbook covering incident classification, escalation paths, and forensic trace retention policies for agentic AI systems.
- Responsible AI (RAI) & Governance (10%)
- Integrate RAI signal capture—fairness, bias detection, explainability, and safety metrics—directly into observability pipelines, making compliance measurable and audit-ready.
- Deliver governance dashboards that surface RAI compliance posture across all active AI agents and LLM deployments, aligned with global regulatory standards.
- Support risk assessments, gap analyses, and governance frameworks with real-time observability insights—enabling proactive risk mitigation rather than reactive audit responses.
- Collaborate with RAI CoE and Legal/Compliance teams to define data retention, consent logging, and model decision traceability standards embedded in the telemetry architecture.
- Quality Engineering for Agentic Solutions — Post Go-Live & Continuous QE (10%)
- Own the Continuous Quality Engineering (CQE) framework for post-production agentic solutions—defining and tracking quality metrics across accuracy, latency, agent success rate, tool-call fidelity, and user outcome measures.
- Build automated quality gates within CI/CD pipelines that leverage observability data to detect regressions, drift, and degradation in agent performance—preventing silent failures in production.
- Instrument and monitor Skill Evaluations (evals) across the Memory, Skills, and MCP harness stack—providing traceability from eval results to production behavior.
- Partner with product and business stakeholders to define SLA-backed quality benchmarks and deliver automated alerting when quality thresholds are breached.
- Drive root-cause analysis for quality failures using distributed trace data, enabling rapid iteration and continuous improvement cycles for agentic solutions.
- Memory, Skills, MCP & Harness Engineering Observability (10%)
- Design and implement observability for the agent memory layer—episodic, semantic, and working memory read/write operations—providing latency, accuracy, and drift monitoring across memory backends.
- Instrument MCP (Model Context Protocol) server interactions, tool registrations, skill invocations, and context injection pipelines with full trace propagation and semantic tagging.
- Own observability for self-evolving harness and reinforcement learning (RL) systems—capturing reward signals, policy update events, environment state transitions, and learning convergence metrics.
- Monitor harness execution fidelity, skill eval pass/fail rates, and regression signals across training, fine-tuning, and inference workflows—feeding data back into the quality engineering loop.
- Data Science Observability & Hardcore Python Engineering (5%)
- Lead a team of senior Python engineers building high-performance, production-grade observability tooling—including custom OTEL exporters, semantic trace enrichers, signal aggregators, and anomaly detection pipelines.
- Apply data science methods—statistical process control, time-series anomaly detection, clustering, and causal inference—to transform raw telemetry into actionable AI operational intelligence.
- Build and maintain Python-native SDKs and libraries that simplify observability onboarding for agent developers across the organization.
- Establish code quality standards, testing frameworks, and peer review practices for the observability engineering team—embedding software craftsmanship into the team culture.
- Agentic Marketplace, Registry & Ecosystem Observability (5%)
- Instrument the Agentic Marketplace and Agent Registry platforms—providing usage telemetry, adoption metrics, capability health scores, and dependency mapping for registered agents and skills.
- Design observability APIs and SDK hooks that allow marketplace-registered agents to self-report health, performance, and behavioral signals into the central observability platform.
- Monitor inter-agent communication patterns across the marketplace ecosystem—identifying latency hotspots, circular dependencies, and protocol mismatches in agent-to-agent (A2A) workflows.
- Deliver a Marketplace Observability Dashboard surfacing agent catalog health, adoption trends, quality scores, and incident history—supporting marketplace governance and curation decisions.
- Integration, Deployment & CI/CD Automation (5%)
- Build and maintain CI/CD pipelines for observability services and agent operations center components, incorporating automated testing, deployment gates, and rollback mechanisms.
- Automate onboarding for new agent use cases using templates, scaffolding, and configuration validation—reducing time-to-observability from weeks to hours.
- Drive infrastructure-as-code (IaC) practices for observability platform components across Azure, AWS, and GCP—ensuring reproducible, version-controlled, and auditable deployments.
- Product Delivery & Stakeholder Collaboration (10%)
- Operate with a product mindset—defining observability platform roadmaps, OKRs, adoption playbooks, and release milestones in partnership with AI platform and business teams.
- Collaborate with transformation teams, enterprise architects, security, and business stakeholders to tailor observability solutions to domain-specific requirements.
- Serve as the technical authority in executive and governance forums—translating complex observability data into business-relevant insights on risk, cost, and AI performance.
- Partner with SRE, AI platform, and product teams to drive standard adoption and reduce integration friction across the agentic AI ecosystem.
- People Leadership & Team Development (5%)
- Build, mentor, and lead a high-performing observability engineering team—spanning Python developers, data scientists, and platform engineers—with talent initially based in India.
- Define career paths, skills development plans, and leveling criteria aligned with PepsiCo job architecture—fostering an inclusive, high-accountability team culture.
- Drive hiring, coaching, performance management, and succession planning across the observability function.
Decision-Making
- Autonomy: High — Owns architecture decisions, platform roadmap, and engineering standards. Strategic alignment sought from AI Solutions Director on enterprise-level commitments.
- Supervision Required: Low to Moderate — Operates independently with periodic alignment reviews. Proactively escalates cross-organizational dependencies and risk trade-offs.
- Role Complexity: Very High — Spans observability, safety/security, RL harnesses, physical AI, multi-modal systems, agent protocols, quality engineering, and marketplace governance simultaneously.
Pay
The expected compensation range for this position is between $123,500 - $206,750. Location, confirmed job-related skills, experience, and education will be considered in setting actual starting salary. Bonus based on performance and eligibility target payout is 15% of annual salary paid out annually.
Benefits
- Paid time off subject to eligibility, including paid parental leave, vacation, sick, and bereavement.
- Comprehensive benefits package to support employees and their families, subject to elections and eligibility:
- Medical, Dental, Vision
- Disability, Health, and Dependent Care Reimbursement Accounts
- Employee Assistance Program (EAP)
- Insurance (Accident, Group