Jobs · Massachusetts

Senior Cloud/Platform Operations Engineer

Bain Capital · Boston, MA · 1 mo ago
Hybrid$155k–$180k/yrFull-time

About the role

We are seeking a Senior Cloud / Platform Operations Engineer to play a key role in the day-to-day operations of our AWS and Kubernetes platforms. This is a hands-on technical role focused on ensuring the reliability, security, and operational excellence of our cloud infrastructure.

Responsibilities

  • Own the day-to-day operations of AWS and Kubernetes environments, ensuring high availability, reliability, and performance.
  • Perform hands-on administration of production Kubernetes clusters and AWS infrastructure.
  • Troubleshoot complex infrastructure, networking, and container platform issues.
  • Execute platform maintenance activities including upgrades, patching, scaling, and lifecycle management.
  • Continuously improve operational processes, automation, and platform stability.
  • Kubernetes Operations:
    • Operate and maintain production Kubernetes clusters.
    • Perform cluster upgrades, patching, and version management.
    • Troubleshoot Kubernetes control plane, worker nodes, networking, storage, ingress, and workload issues.
    • Optimize cluster performance, resource utilization, and resilience.
    • Support containerized application deployments and resolve runtime issues.
    • Implement Kubernetes operational best practices around security, reliability, and scalability.
  • AWS Operations & Governance:
    • Support AWS account administration, IAM, networking, and infrastructure management.
    • Implement and maintain AWS governance guardrails, security controls, and access policies.
    • Ensure compliance with organizational standards for cloud infrastructure.
    • Assist with cloud cost optimization and resource management.
    • Partner with security teams to remediate infrastructure risks and vulnerabilities.
  • Incident Response & Operational Support:
    • Participate in production incident response and on-call rotations.
    • Troubleshoot and resolve complex production issues across AWS and Kubernetes environments.
    • Perform root cause analysis and implement corrective actions.
    • Support a ticket-driven operational model with a strong focus on responsiveness and customer service.
    • Document operational procedures and contribute to knowledge sharing across the team.
  • Monitoring & Reliability:
    • Maintain monitoring, logging, and alerting for cloud infrastructure and Kubernetes platforms.
    • Proactively identify reliability risks before they impact users.
    • Improve observability and operational visibility across the platform.
    • Drive continuous improvements to platform uptime and operational efficiency.
  • Cross-Functional Collaboration:
    • Partner with engineering teams to troubleshoot application deployment issues.
    • Support internal users by resolving infrastructure requests and platform-related issues.
    • Collaborate with security, networking, and development teams to deliver reliable cloud services.
    • Contribute technical expertise during infrastructure planning and operational improvements.

Qualifications

  • 5+ years of experience supporting cloud infrastructure in production environments.
  • Strong hands-on experience administering production Kubernetes environments, including: Cluster installation and lifecycle management, Upgrades and patching, Networking (CNI), Ingress controllers, Persistent storage, RBAC and security, Troubleshooting workloads and cluster health.
  • Strong hands-on experience with AWS services including: IAM, VPC, EC2, EKS, Route 53, CloudWatch, Load Balancers, EFS.
  • Experience implementing AWS governance, IAM policies, and security best practices.
  • Strong Linux systems administration skills.
  • Experience with Infrastructure as Code (Terraform and Ansible preferred).
  • Experience with CI/CD pipelines and deployment automation.
  • Experience with monitoring and observability tools such as Datadog or similar.
  • Experience supporting production incidents and participating in on-call rotations.
  • Strong troubleshooting and root cause analysis skills.
  • Excellent communication skills with the ability to work directly with technical and non-technical stakeholders.

Preferred Qualifications

  • Experience operating Kubernetes platforms at enterprise scale.
  • Experience with GitOps tools such as Argo CD.
  • Experience with policy enforcement tools such as Kyverno.
  • Experience with container security platforms and vulnerability management.
  • AWS Solutions Architect, AWS SysOps Administrator, or Certified Kubernetes Administrator (CKA) certification.
  • Experience in regulated or security-conscious environments.

Similar jobs