Senior / Staff Cloud Platform Engineer (ML)
About the role
Calico seeks a Senior / Staff Cloud Engineer to lead the execution and operation of our next-generation machine learning (ML) platform. As the technical authority for our Google Cloud Platform (GCP) environment, you will build infrastructure-as-Code (IaC), design secure perimeters, automate identity management, and migrate existing workflows into a unified, compliant infrastructure. This role is ideal for someone passionate about DevSecOps, IaC, and building "paved paths" that empower scientists while enforcing rigorous security. You will serve as the critical bridge enabling ML Systems and Research engineers to train and deploy ML models for frontier BioML research.
No biology or life sciences background is required for this role.
Responsibilities
- Foundational Impact:
- Deliver our next-generation ML platform
- Lead the consolidation and modernization of our compute environments to accelerate the company's R&D cycle
- GCP Foundation:
- Lead the centralization of GCP projects into a unified, compliant landing zone
- Design secure perimeters to protect sensitive biological data
- Build automated guardrails using Terraform
- GKE Cluster Administration:
- Design, deploy, and maintain robust GKE clusters tailored for distributed ML training and batch processing workloads
- Troubleshoot node-level GPU/TPU issues
- Manage scheduling add-ons and optimize cluster autoscaling
- Developer Experience & Onboarding:
- Design intuitive onboarding processes and "paved paths" for scientists to seamlessly onboard onto the new platform
- FinOps:
- Manage org-level GPU/TPU compute capacity
- Implement centralized billing and cost reporting
- Build comprehensive monitoring dashboards to track cluster utilization and optimize R&D expenditure
Requirements
- BS/MS with 7+ years or PhD with 4+ years of relevant software engineering, DevSecOps, or Cloud architecture experience
- Strong programming skills in Python
- Strong hands-on experience with Google Cloud Platform (GCP) core services
- Experience with Kubernetes cluster administration
- Experience with Terraform, building CI/CD pipelines, and containerization (Docker)
- Must be willing to work onsite at least four days a week
Nice to Have
- A strong intellectual curiosity for life sciences
- Experience managing Kubernetes for HPC or ML workloads
- Experience supporting Machine Learning orchestration and queueing frameworks (e.g., Kueue, Ray)
- Experience reading and debugging ML Python code
- Experience implementing compliant cloud architectures (e.g., SOC2, HIPAA, or similar frameworks)
- Familiarity with Golang
- Strong communication skills with a proven ability to guide users through migration
Pay
The estimated base salary range for this role is $220,000 - $290,000. Actual pay will be based on experience and qualifications. This position is also eligible for two annual cash bonuses.