Jobs · Engineering · Texas

Site Reliability Engineering (SRE) Manager

IBM · Austin, TX · 1 mo ago
On-siteEngineeringFull-time

About the Role

We are seeking a Site Reliability Engineer (SRE) Manager to lead a high-performing SRE team within the IBM CISO Platform team. You will drive team execution, partner with cross-functional teams, and ensure IBM’s internal security platforms maintain world-class performance, resilience, and compliance. This role values innovative thinkers, strong leaders, and individuals passionate about continuous improvement in a dynamic security environment.

Responsibilities

  • Lead and mentor a team of Site Reliability Engineers, providing coaching, career development, and performance management.
  • Oversee the implementation and automation of operational processes, infrastructure, monitoring, incident response, and runbooks.
  • Own end-to-end service reliability, including SLI/SLOs, capacity planning, performance optimization, and operational health.
  • Ensure platforms meet IBM CISO and enterprise security standards, regulatory requirements, and risk policies.
  • Communicate strategy, risks, operational status, and metrics to leadership and stakeholders.
  • Influence technology roadmaps and operational readiness for new internal solutions.
  • Foster a high-performing engineering culture centered around accountability, innovation, and continuous improvement.
  • Align team objectives with the strategic direction of the IBM CISO organization and broader Enterprise & Technology Services.
  • Plan staffing, manage workload distribution, and ensure on-call readiness and 24/7 service support coverage.
  • Develop component-level software solutions, ensuring implemented solutions are unit tested and ready for integration.
  • Contribute to the automated CI/CD pipeline for seamless integration and delivery.
  • Debug and resolve customer-reported problems through code fixes and collaboration with stakeholders.
  • Deliver high-quality offerings using leading-edge or proven technologies in an Agile environment.
  • Balance support of current systems while leading modernization and future-state design.
  • Handle critical issues outside of business hours and manage Release/Change Management processes.

Requirements

  • Proven experience managing or leading engineering, SRE, DevOps, or operations teams.
  • Strong background in delivering reliable, highly available services.
  • Deep understanding of security, compliance, and risk management frameworks.
  • Demonstrated success driving automation of infrastructure, monitoring, and operational tasks.
  • Excellent written and verbal communication skills with the ability to influence and drive alignment across teams.
  • Experience with Release/Change Management processes.
  • Ability to handle critical issues outside of business hours.

Preferred Qualifications

  • Experience with Kubernetes, OpenShift, or similar container orchestration platforms.
  • Experience building or operating Cloud-native environments (AWS, Azure, GCP, IBM Cloud), Hybrid Cloud, and on-prem infrastructure.
  • Familiarity with observability tools.
  • Understanding of networking fundamentals and modern networking architectures.
  • Knowledge of Infrastructure as Code (Terraform, Ansible, etc.).
  • Exposure to Agile methodologies (Jira, Kanban, Scrum, etc.).
  • Working knowledge or scripting/programming languages (e.g., Python).
  • Professional Cloud and/or Security certifications (AWS, CISSP, etc.).

Similar jobs

Site Reliability Engineer (SRE)

Bright Vision TechnologiesFramingham, MA· 4 days ago
RemoteEngineering$100k–$180k/yrapply on brightvisiontechnologies.applytojob.com