Site Reliability Engineering (SRE) Manager
IBM · Austin, TX · 1 mo ago
On-siteEngineeringFull-time
About the Role
We are seeking a Site Reliability Engineer (SRE) Manager to lead a high-performing SRE team within the IBM CISO Platform team. You will drive team execution, partner with cross-functional teams, and ensure IBM’s internal security platforms maintain world-class performance, resilience, and compliance. This role values innovative thinkers, strong leaders, and individuals passionate about continuous improvement in a dynamic security environment.
Responsibilities
- Lead and mentor a team of Site Reliability Engineers, providing coaching, career development, and performance management.
- Oversee the implementation and automation of operational processes, infrastructure, monitoring, incident response, and runbooks.
- Own end-to-end service reliability, including SLI/SLOs, capacity planning, performance optimization, and operational health.
- Ensure platforms meet IBM CISO and enterprise security standards, regulatory requirements, and risk policies.
- Communicate strategy, risks, operational status, and metrics to leadership and stakeholders.
- Influence technology roadmaps and operational readiness for new internal solutions.
- Foster a high-performing engineering culture centered around accountability, innovation, and continuous improvement.
- Align team objectives with the strategic direction of the IBM CISO organization and broader Enterprise & Technology Services.
- Plan staffing, manage workload distribution, and ensure on-call readiness and 24/7 service support coverage.
- Develop component-level software solutions, ensuring implemented solutions are unit tested and ready for integration.
- Contribute to the automated CI/CD pipeline for seamless integration and delivery.
- Debug and resolve customer-reported problems through code fixes and collaboration with stakeholders.
- Deliver high-quality offerings using leading-edge or proven technologies in an Agile environment.
- Balance support of current systems while leading modernization and future-state design.
- Handle critical issues outside of business hours and manage Release/Change Management processes.
Requirements
- Proven experience managing or leading engineering, SRE, DevOps, or operations teams.
- Strong background in delivering reliable, highly available services.
- Deep understanding of security, compliance, and risk management frameworks.
- Demonstrated success driving automation of infrastructure, monitoring, and operational tasks.
- Excellent written and verbal communication skills with the ability to influence and drive alignment across teams.
- Experience with Release/Change Management processes.
- Ability to handle critical issues outside of business hours.
Preferred Qualifications
- Experience with Kubernetes, OpenShift, or similar container orchestration platforms.
- Experience building or operating Cloud-native environments (AWS, Azure, GCP, IBM Cloud), Hybrid Cloud, and on-prem infrastructure.
- Familiarity with observability tools.
- Understanding of networking fundamentals and modern networking architectures.
- Knowledge of Infrastructure as Code (Terraform, Ansible, etc.).
- Exposure to Agile methodologies (Jira, Kanban, Scrum, etc.).
- Working knowledge or scripting/programming languages (e.g., Python).
- Professional Cloud and/or Security certifications (AWS, CISSP, etc.).