Cloud/Platform Operations Manager
Bain Capital · Boston, MA · 2 wk ago
Hybrid$180k–$210k/yrFull-time
About the role
We are seeking a hands-on, operationally focused Cloud / Platform Operations Manager to lead the day-to-day management of our AWS and Kubernetes environments.
Responsibilities
- Own the day-to-day operations of AWS and Kubernetes environments, ensuring reliability, availability, and performance.
- Provide technical leadership and hands-on support for complex Kubernetes and cloud infrastructure issues.
- Lead a ticket-driven support model, prioritizing and ensuring timely fulfillment of internal user requests.
- Establish and enforce SLAs/SLOs for platform support and operational responsiveness.
- Drive a culture of execution, accountability, and customer service within the team.
- Define, implement, and enforce AWS guardrails, IAM policies, and governance frameworks.
- Ensure proper account structure, access controls, cost controls, and compliance standards.
- Partner with security teams to enforce best practices and reduce risk across the platform.
- Continuously audit and improve cloud policy adherence.
- Own the operational health and lifecycle management of production Kubernetes clusters.
- Maintain and troubleshoot Kubernetes control planes, worker nodes, networking, storage, ingress, and containerized workloads.
- Perform hands-on cluster administration, upgrades, patching, capacity planning, and performance tuning.
- Ensure robust monitoring, logging, alerting, and incident response for Kubernetes environments.
- Drive operational best practices for cluster reliability, security, and resiliency.
- Partner with application teams to troubleshoot deployment, scaling, networking, and workload issues.
- Participate in a rotating after-hours on-call schedule to provide operational support for critical production incidents, ensuring timely response, issue resolution, and service continuity.
- Act as the primary escalation point for AWS and Kubernetes operational issues.
- Lead technical troubleshooting during production incidents and guide the team through resolution.
- Handle and resolve escalations effectively before they reach senior leadership.
- Build strong communication channels with stakeholders during incidents.
- Lead, coach, and develop a team of cloud/platform engineers with a strong operational mindset.
- Mentor engineers on Kubernetes operations, AWS best practices, and operational excellence.
- Set clear expectations around ownership, responsiveness, and execution.
- Manage performance, provide feedback, and address gaps in delivery or behavior.
- Foster a culture of accountability, urgency, and continuous improvement.
- Partner with engineering, product, and business teams to support platform needs and unblock users.
- Balance operational workload with longer-term improvements without compromising service delivery.
- Act as a bridge between users and platform capabilities, ensuring needs are met efficiently.
Qualifications
- 8+ years of experience in cloud infrastructure or platform operations, with at least 2 years in a leadership role.
- Strong hands-on experience administering and supporting production Kubernetes environments, including cluster operations, troubleshooting, upgrades, networking, storage, ingress, and workload management.
- Demonstrated ability to independently diagnose and resolve complex Kubernetes platform issues in production.
- Strong hands-on experience with AWS, including IAM, networking, account structure, and governance controls.
- Proven experience managing AWS guardrails, policies, and access controls at scale.
- Experience with Kubernetes observability and operational tooling (such as Prometheus, Grafana, CloudWatch, Fluent Bit, or similar).
- Deep operational experience running and supporting Kubernetes environments in production.
- Experience leading ticket-based support or operations teams with high request volume.
- Strong incident management and escalation handling experience.
- Demonstrated ability to lead teams focused on execution, reliability, and service delivery while remaining technically engaged.
- Excellent communication and stakeholder management skills.