Senior DevOps Engineer (EKS/Kubernetes)
Caris Life Sciences · United States · Today
RemoteRemoteEngineering$122k–$148k/yrFull-time
Job Responsibilities
- Design, deploy, and maintain Linux infrastructure in on-premises and cloud environments.
- Automate infrastructure provisioning and configuration using tools such as Terraform, Ansible, or CloudFormation.
- Manage and optimize AWS environments with a focus on performance, scalability, security, and cost efficiency.
- Implement and maintain monitoring, logging, and alerting solutions (e.g., Datadog, Prometheus, Grafana, ELK, CloudWatch).
- Architect, deploy, and operate production Kubernetes/AWS EKS clusters, including node group strategy, cluster upgrades, multi-tenant workload isolation, and cross-region disaster recovery (DR) architecture and build outs.
- Define and lead cluster upgrade, security hardening, and disaster recovery strategies for production Kubernetes/AWS EKS environments at scale, while serving as a senior technical resource for complex production incidents.
- Manage Kubernetes networking, including VPC CNI configuration and ingress controllers (ALB/NGINX/Traefik).
- Implement IAM Roles for pod security standards, and network policies to secure EKS workloads.
- Configure and tune cluster autoscaling (Cluster Autoscaler or Karpenter) and workload autoscaling (HPA/VPA) to optimize performance and cost.
- Build and maintain Helm charts and GitOps-based deployment pipelines (e.g., ArgoCD, Flux) for Kubernetes workloads.
- Manage Docker container builds and registries in support of EKS-based application deployment.
- Deploy, scale, and maintain GitLab Runners (including Kubernetes executor runners on EKS) to support CI/CD pipeline throughput and reliability.
- Support and help operate database platforms on AWS RDS (MySQL, PostgreSQL), collaborating with data owners on performance and reliability.
- Ensure systems meet security and compliance requirements, including SOX and SOC 2 initiatives.
- Execute and maintain Linux patching strategies, addressing security updates and CVEs in a timely manner.
- Support and help operate incident response, root cause analysis, and recovery efforts.
- Collaborate with development, QA, and cross-functional teams to improve reliability, release processes, and operational standards.
- Participate in on-call rotations and provide after-hours support as required.
Required Qualifications
- Bachelor’s degree in computer science, Information Technology or related field.
- 8+ years of experience in Linux Systems Administration, DevOps, or Site Reliability Engineering roles.
- 5+ years of experience with AWS services, including EC2, VPC, IAM, RDS, S3, and CloudWatch.
- 5+ years of hands-on experience designing and operating production workloads on Kubernetes/AWS EKS, including cluster upgrades, networking, and autoscaling.
- Proficiency in scripting and automation using Python and Bash.
- Strong hands-on experience with Infrastructure as Code using Terraform, and with CI/CD pipelines (GitLab CI/CD), including running CI/CD workloads on Kubernetes/EKS.
- Proficiency with Docker, Helm, and Kubernetes troubleshooting in a production environment.
- Solid understanding of networking fundamentals and cloud security best practices.
- CKA (Certified Kubernetes Administrator) certification expected or actively in progress.
Preferred Qualifications
- CKAD and AWS certifications (e.g., AWS Certified DevOps Engineer, Solutions Architect) a plus.
- Experience with Karpenter, Kyverno, OPA/Gatekeeper, Falco, and multi-cluster/multitenant EKS environments.
- Experience using AI and automation tools including Claude, Cursor, and OpenAI to streamline DevOps workflows through AI-assisted CI/CD, self-healing operations, and automated incident response.
- Experience with microservices, serverless architectures, and DevSecOps practices.