Principal Customer Experience Program Manager
About the Role
We are hiring a Principal Customer Experience Program Manager to lead two interconnected workstreams under Advanced Cloud Engineering & Supportability (ACES) Sovereign & Government:
- Gov Customer Resiliency (60%): Build and operate a new Gov Customer Resiliency function from scratch, starting with a named high-profile Government customer and scaling to a portfolio of 3-5 top Gov/Azure Engineering Direct customers. Own the full resiliency lifecycle: proactive detection and monitoring, incident and crisis management coordination, post-incident RCA and problem management, architecture and DR guidance, and parity closure between Government and Commercial cloud environments. This is a build + run role—define the operating model, codify playbooks, and scale the function across regulated environments.
- Sovereign Cloud Operations & Readiness (40%): Drive support readiness, operational maturity, and customer experience strategy across Microsoft's Sovereign Cloud portfolio (Bleu, Delos, Merlion). Design escalation pathways, incident handling standards, compliance-aligned processes, and readiness frameworks for new Sovereign cloud launches. Partner with Sovereign delivery leadership, Azure engineering, and regional National Cloud Operating Entity (NCOE) partners to ensure support readiness and exceptional customer outcomes from Day 1.
This role sits at the intersection of two strategic ACES investments: bringing proactive reliability engineering in-house for Government customers and ensuring Sovereign clouds are support-ready from launch. You will shape how Microsoft supports its highest-trust customers and define a practice that will serve as a model for the broader organization.
Responsibilities
- Gov Customer Resiliency (60%):
- Stand up a new proactive resiliency function for Government cloud customers: define charter, build playbooks, establish operating cadences, and own the end-to-end engagement model.
- Own the full resiliency lifecycle: proactive detection and monitoring, incident and crisis coordination, post-incident root cause analysis, and architecture/DR guidance.
- Drive Gov-vs-Commercial parity closure across monitoring, tooling, incident response, and remediation maturity.
- Lead resiliency and reliability workshops and customer conversations, including Field enablement teams.
- Scale the resiliency model from a single anchor customer to a portfolio of 3-5 top Government customers using a repeatable, metrics-driven playbook.
- Develop and deliver internal enablement content (training materials, case studies, learning sessions) to embed resiliency practices across Gov Support delivery teams.
- Define and report on success metrics: mean time to detect, time to engage, incident recurrence, proactive detection rates, and customer confidence.
- Leverage telemetry, monitoring data, and trend analysis to proactively identify and address emerging risks before they become customer-reported incidents.
- Partner with reliability engineering, product teams, and delivery leadership to ensure resiliency insights feed into upstream engineering actions and product improvements.
- Sovereign Cloud Operations & Readiness (40%):
- Drive end-to-end support readiness (people, process, technology) for Microsoft's Sovereign Cloud portfolio across multiple regions and future launches.
- Design escalation pathways, incident handling standards, and compliance-aligned operational processes for Sovereign environments.
- Own readiness frameworks for new Sovereign cloud launches; influence design decisions upstream to prevent customer impact.
- Lead operational reporting and insights; translate data into risk assessments and executive-ready recommendations.
- Represent Sovereign and Government customer needs in cross-org forums, influencing priorities and investments to strengthen long-term customer trust.
- Cross-Cutting:
- Leverage AI, automation, and data-driven insights to proactively identify gaps, reduce risk, and improve customer experience at scale.
- Extend the Gov Resiliency playbook to Sovereign clouds as they mature, building a unified approach across regulated environments.
- Drive alignment across geographically distributed teams and operating partners spanning multiple countries and time zones.
Requirements
- Required Qualifications:
- Bachelor's Degree in Computer Science, Engineering, Data Science, Math, Business, or related field AND 6+ years' experience in engineering, product/technical program management, data analysis, or product development OR equivalent experience.
- US Citizenship: This position requires verification of US citizenship due to citizenship-based legal restrictions for supporting US federal, state, and/or local government agency customers.
- Microsoft Cloud Background Check: Required upon hire/transfer and every two years thereafter.
- Preferred Qualifications:
- Master's Degree in Computer Science, Engineering, Data Science, Math, Business, or related field AND 8+ years' experience OR Bachelor's Degree AND 12+ years' experience in engineering, product/technical program management, data analysis, or product development.
- Experience in CRE, SRE, ACE, or operational reliability roles within a cloud hyperscaler environment.
- Hands-on experience with resiliency tooling, platform monitoring, and incident management systems.
- Deep knowledge of Sovereign compliance, data residency, and geo-centric architecture models (e.g., EU Data Boundary, government cloud isolation requirements).
- Track record of executive-level customer engagement, including leading confidence calls, MBRs, and exec-level progress reviews with enterprise customers.
- Demonstrated experience in customer-facing resiliency, reliability engineering, or incident management roles, including proactive detection, crisis coordination, or post-incident program management.
- Customer and Field-facing experience driving deep technical and architecture conversations, including resiliency workshops.
- Experience working with government agencies, sovereign entities, or regulated industries, with a strong understanding of their missions, operating models, compliance requirements, and IT environments.
- Strong understanding of Azure services and cloud technologies, including monitoring, diagnostics, incident response tooling, and infrastructure architecture.
- Proven ability to build new functions or programs from scratch, defining charter, playbooks, metrics, operating cadences, and scaling across customers.
- Exceptional cross-org stakeholder management skills; ability to drive alignment across engineering, support delivery, product teams, and customer-facing partners without direct authority.
- Experience working effectively across multiple geographies, cultures, and organizational boundaries.
Location & Schedule
- Based in the United States. Atlanta, GA strongly preferred (proximity to Gov customers and CRE team members).
- East Coast candidates preferred for European timezone overlap (Sovereign clouds in France, Germany, Singapore). West Coast candidates must support early-morning collaboration with Europe.
- Travel: In-office at least three days per week (subject to local policy).
Pay
The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. A different range applies to specific work locations, including the San Francisco Bay area and New York City metropolitan area, where the base pay range is USD $188,000 - $304,200 per year. Certain roles may be eligible for benefits and other compensation. Additional pay information can be found here.