Senior Platform Engineer II, Compute Services
CoreWeave · Sunnyvale, CA · 1 wk ago
Engineering$165k–$242k/yrFull-time
About the role
We are seeking a Senior Platform Engineer to join our Kubernetes Infrastructure team. This role involves administering our critical multi-tenant Kubernetes platforms and collaborating with development teams to establish proper deployment architectures.
Responsibilities
- Champion reliability initiatives for Kubernetes application deployments: Advocate for best practices to ensure high availability, scalability, and resilience of applications in Kubernetes, focusing on robust testing, secure pipelines, and efficient resource use.
- Administer multi-tenant Kubernetes platforms: Manage complex multi-tenant Kubernetes clusters, configuring access, quotas, and security for isolation and optimal resource allocation while upholding SLAs.
- Perform lifecycle and day 2 operations on clusters: Execute Kubernetes cluster lifecycle, including provisioning, patching, monitoring, backup, disaster recovery, and troubleshooting.
- Deep dive into reliability issues: Conduct in-depth analysis and root cause identification for complex reliability incidents in Kubernetes, utilizing advanced debugging and monitoring tools to propose preventative measures.
- Respond to critical alerts and incidents outside business hours, providing timely resolution to minimize disruptions, collaborating with teams, and communicating clearly.
Requirements
- Bachelor's in CS, Engineering, or related field, or equivalent experience preferred.
- CKA or similar certifications is highly desired.
- 5+ years administering multi-tenant SaaS Kubernetes (EKS, AKS, GKS).
- Strong Gitops/DevOps with ArgoCD or similar helm chart management.
- Proven Docker and containerization experience.
- Strong Linux OS experience.
- Proficient in Go.
- Excellent problem-solving, debugging, and analytical skills.
- Strong communication and collaboration.
Qualifications
- Preferred Master's degree in Computer Science, Engineering, or a related field.
- Experience with performance profiling and optimization of distributed systems.
- Knowledge of network protocols and distributed consensus algorithms.
Skills
- Experience with Kubernetes administration and management.
- Strong understanding of cloud-native technologies and infrastructure.
- Ability to troubleshoot and resolve complex technical issues.
- Effective communication and collaboration skills.
Benefits
- The base salary range for this role is $165,000 to $242,000.
- A comprehensive benefits program including medical, dental, and vision insurance, company-paid life insurance, short and long-term disability insurance, flexible spending account, health savings account, tuition reimbursement, mental wellness benefits, employee stock purchase program, and more.