Jobs · Engineering · New York

Staff Site Reliability Engineer - Kubernetes

Okta · New York, NY · 2 days ago
HybridEngineering$194k/yrFull-time

Key Responsibilities

  • Kubernetes Platform Creation: Design, implement, and maintain highly available, scalable, and fault-tolerant Kubernetes platforms. Ensure clusters are optimized for production workloads, providing high resilience and operational efficiency.
  • AWS Infrastructure Management: Build, manage, and optimize AWS cloud infrastructure, including EKS, ECS, S3, VPCs, RDS, IAM, and more. Implement best practices for cost management, scaling, and security within AWS.
  • Helm Management: Utilize Helm to automate and streamline the deployment of applications and services to Kubernetes clusters. Create, maintain, and manage Helm charts for production-ready deployments.
  • Karpenter Implementation: Implement and manage Karpenter to dynamically scale Kubernetes clusters in response to workload demands. Optimize resource usage.
  • Istio Service Mesh Management: Configure and manage Istio to provide service-to-service communication, security, and observability within the Kubernetes clusters. Enable fine-grained traffic management, service discovery, and policy enforcement.
  • Platform Automation & Scaling: Automate the deployment, scaling, and management of infrastructure and applications. Work with CI/CD pipelines to ensure a seamless flow from development to production with minimal downtime.
  • Incident Management & Troubleshooting: Respond to incidents, troubleshoot, and resolve system issues related to performance, availability, and security in a timely and effective manner.

Required Qualifications

  • 4+ years of experience with Kubernetes/Helm;
  • 4+ years of Experience with Terraform;
  • 5+ years of Experience with AWS;
  • Proven experience with AWS (EC2, RDS, S3, CloudFormation, IAM, etc.) and solid understanding of cloud-native architectures;
  • Strong expertise in Kubernetes platform creation, management, and optimisation (e.g., setting up highly available clusters, networking, and storage);
  • Hands-on experience with Helm for Kubernetes application deployment and management;
  • Practical experience with Karpenter for dynamic scaling of Kubernetes clusters and optimising resource usage.
  • Expertise in managing and securing Istio for service mesh, including traffic management, security, and observability features;
  • Proficiency in CI/CD pipelines and automation tools (e.g., Jenkins, GitLab, CircleCI, Terraform, Ansible, Spinnaker).
  • Strong scripting and automation skills in Python, Bash, or Go for infrastructure management and platform automation.
  • Experience with monitoring, logging, and alerting tools such as Prometheus, Grafana, CloudWatch, and ELK Stack.

Preferred Qualifications

  • Understanding of security best practices for cloud platforms and Kubernetes (e.g., role-based access control (RBAC), encryption, and compliance frameworks).
  • Familiarity with Docker and containerization principles.
  • Bachelor’s degree in Computer Science, Engineering, or related field (or equivalent professional experience).
  • Certifications (Preferred): CKA (Certified Kubernetes Administrator), CKAD (Certified Kubernetes Application Developer), or AWS Certified DevOps Engineer are highly desirable.

Similar jobs

Staff Software Engineer

Bright Vision TechnologiesSpringfield, Illinois, United States· 1 mo ago
RemoteEngineeringapply on brightvisiontechnologies.applytojob.com

Staff Software Engineer

PayNearMeUnited States· 2 mo ago
RemoteEngineering$210k–$255k/yrapply on job-boards.greenhouse.io