Senior Cloud/Platform Operations Engineer
Bain Capital · Boston, MA · 1 mo ago
Hybrid$155k–$180k/yrFull-time
About the role
We are seeking a Senior Cloud / Platform Operations Engineer to play a key role in the day-to-day operations of our AWS and Kubernetes platforms. This is a hands-on technical role focused on ensuring the reliability, security, and operational excellence of our cloud infrastructure.
Responsibilities
- Own the day-to-day operations of AWS and Kubernetes environments, ensuring high availability, reliability, and performance.
- Perform hands-on administration of production Kubernetes clusters and AWS infrastructure.
- Troubleshoot complex infrastructure, networking, and container platform issues.
- Execute platform maintenance activities including upgrades, patching, scaling, and lifecycle management.
- Continuously improve operational processes, automation, and platform stability.
- Kubernetes Operations:
- Operate and maintain production Kubernetes clusters.
- Perform cluster upgrades, patching, and version management.
- Troubleshoot Kubernetes control plane, worker nodes, networking, storage, ingress, and workload issues.
- Optimize cluster performance, resource utilization, and resilience.
- Support containerized application deployments and resolve runtime issues.
- Implement Kubernetes operational best practices around security, reliability, and scalability.
- AWS Operations & Governance:
- Support AWS account administration, IAM, networking, and infrastructure management.
- Implement and maintain AWS governance guardrails, security controls, and access policies.
- Ensure compliance with organizational standards for cloud infrastructure.
- Assist with cloud cost optimization and resource management.
- Partner with security teams to remediate infrastructure risks and vulnerabilities.
- Incident Response & Operational Support:
- Participate in production incident response and on-call rotations.
- Troubleshoot and resolve complex production issues across AWS and Kubernetes environments.
- Perform root cause analysis and implement corrective actions.
- Support a ticket-driven operational model with a strong focus on responsiveness and customer service.
- Document operational procedures and contribute to knowledge sharing across the team.
- Monitoring & Reliability:
- Maintain monitoring, logging, and alerting for cloud infrastructure and Kubernetes platforms.
- Proactively identify reliability risks before they impact users.
- Improve observability and operational visibility across the platform.
- Drive continuous improvements to platform uptime and operational efficiency.
- Cross-Functional Collaboration:
- Partner with engineering teams to troubleshoot application deployment issues.
- Support internal users by resolving infrastructure requests and platform-related issues.
- Collaborate with security, networking, and development teams to deliver reliable cloud services.
- Contribute technical expertise during infrastructure planning and operational improvements.
Qualifications
- 5+ years of experience supporting cloud infrastructure in production environments.
- Strong hands-on experience administering production Kubernetes environments, including: Cluster installation and lifecycle management, Upgrades and patching, Networking (CNI), Ingress controllers, Persistent storage, RBAC and security, Troubleshooting workloads and cluster health.
- Strong hands-on experience with AWS services including: IAM, VPC, EC2, EKS, Route 53, CloudWatch, Load Balancers, EFS.
- Experience implementing AWS governance, IAM policies, and security best practices.
- Strong Linux systems administration skills.
- Experience with Infrastructure as Code (Terraform and Ansible preferred).
- Experience with CI/CD pipelines and deployment automation.
- Experience with monitoring and observability tools such as Datadog or similar.
- Experience supporting production incidents and participating in on-call rotations.
- Strong troubleshooting and root cause analysis skills.
- Excellent communication skills with the ability to work directly with technical and non-technical stakeholders.
Preferred Qualifications
- Experience operating Kubernetes platforms at enterprise scale.
- Experience with GitOps tools such as Argo CD.
- Experience with policy enforcement tools such as Kyverno.
- Experience with container security platforms and vulnerability management.
- AWS Solutions Architect, AWS SysOps Administrator, or Certified Kubernetes Administrator (CKA) certification.
- Experience in regulated or security-conscious environments.