Senior Software Engineering Manager
About the role
We are seeking an experienced engineering manager to lead a team of Site Reliability Engineers responsible for the availability, resiliency, security, and operational excellence of large-scale ad serving systems across hybrid on-premises and Azure environments. In this role, you will guide the design and operation of observability, automation, incident response, scaling, and failover capabilities; drive continuous improvement through blameless postmortems and infrastructure-as-code practices; and partner with machine learning, platform, and engineering teams to improve developer experience and accelerate reliable research-to-production workflows.
Starting January 26, 2026, Microsoft AI (MAI) employees who live within a 50-mile commute of a designated Microsoft office in the U.S. or 25-mile commute of a non-U.S., country-specific location are expected to work from the office at least four days per week. This expectation is subject to local law and may vary by jurisdiction.
Responsibilities
- Team leadership: Lead a team of experienced SREs to ensure uptime, resiliency and fault tolerance of hybrid on-prem and Azure ad serving systems
- Observability: Design and help maintain monitoring, alerting, and logging systems to provide real-time visibility into platform performance
- Automation & Tooling: Lead building of automation for deployments, incident response, scaling, and failover in hybrid cloud/on-prem environments
- Incident Management: Lead on-call rotations, troubleshoot production issues, conduct blameless postmortems, and drive continuous improvements
- Security & Compliance: Ensure data privacy, compliance, and secure operations across serving environments
- Collaboration: Partner with ML engineers and platform teams to improve developer experience and accelerate research-to-production workflows
Requirements
- Bachelor's Degree in Computer Science or related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience
- Ability to meet Microsoft, customer and/or government security screening requirements, including the Microsoft Cloud Background Check (required upon hire/transfer and every two years thereafter)
Qualifications
- Master's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR
- Bachelor's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience
- 4+ years managing Linux and open source systems
- 4+ years people management experience
- 6+ years experience in monitoring & observability tools (Grafana, Datadog, OpenTelemetry, etc.)
- Experience migrating complex systems from on-prem to public cloud
- Experience using Infrastructure-as-code (e.g. Terraform, Bicep) to manage complex systems
- Knowledge of CI/CD pipelines
- Solid knowledge of distributed systems, networking, and storage
Pay
The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year. Certain roles may be eligible for benefits and other compensation.