Senior Site Reliability Engineer – Compute Platforms
About the role
We are seeking a highly experienced Senior Site Reliability Engineer – Compute Platforms to design, implement, and support Kubernetes on baremetal and hypervisor platforms in a private cloud environment. This role is responsible for the architecture, design, and standardization of enterprise compute and hypervisor environments spanning bare metal infrastructure, operating systems, hypervisors, private cloud orchestration, and Kubernetes using Infrastructure-as-Code and GitOps practices. This is a deeply technical role requiring expert-level understanding of compute hardware management, Kubernetes, OpenStack, hypervisors, and extensive working knowledge of Linux Operating systems.
You will collaborate with platform and SRE teams to maintain secure, performant, and multi-tenant-isolated services that serve high-throughput, mission-critical applications.
Responsibilities
- Lead the architecture and design of enterprise compute and hypervisor platform solutions across hardware, OS, virtualization, cloud orchestration, and container orchestration layers
- Define standards and automation frameworks for bare metal provisioning and lifecycle management
- Design and implement Bare Metal as a Service (BMaaS) capabilities for scalable infrastructure consumption
- Architect and design Kubernetes platforms on bare metal with QoS and Affinity (ArgoCD)
- Architect and validate automated deployments of operating systems and hypervisors including Ubuntu and Harvester
- Design and maintain PXE-based provisioning environments leveraging Redfish APIs for large-scale server deployments
- Develop Infrastructure-as-Code using Ansible, Terraform, Helm, and Git, with Python/Bash automation
- Implement CI/CD pipelines for infrastructure updates, patching, upgrades, testing, and rollback
- Design automated workflows for server build, firmware lifecycle management, patching, and hardware validation
- Evaluate and standardize enterprise hardware platforms to meet performance, scalability, and reliability requirements
- Produce detailed high-level and low-level design documentation, build guides, and operational handoff materials
- Perform deep troubleshooting across storage, Kubernetes, hypervisors, networking, and Linux systems
- Partner with operations, network, storage, and platform teams to ensure designs are supportable and production-ready
- Participate in on-call escalation support for complex platform-related issues
- Collaborate globally on change management, documentation, and operational best practices
Requirements
- 6+ years of experience in infrastructure engineering, platform engineering, or DevOps with a strong focus on Compute system design
- Proven experience designing and automating bare metal compute environments at scale
- Strong hands-on experience with PXE boot, network-based OS provisioning, and automated server imaging
- Experience implementing or supporting Bare Metal as a Service (BMaaS) platforms
- Practical experience using Redfish APIs for hardware provisioning, power management, and remote lifecycle operations
- Deep expertise with Ubuntu Linux in enterprise environments
- Strong hands-on experience with KVM hypervisors (Suse Harvester, OpenStack)
- Experience designing and deploying production-grade Kubernetes clusters
- Strong background with enterprise compute hardware platforms, including Cisco UCS, Dell PowerEdge, Supermicro systems & HPE
- Proficiency with Infrastructure as Code tools (e.g., Terraform, Ansible, or similar)
- Experience building or supporting CI/CD pipelines for infrastructure and platform automation
- Strong scripting skills in Python, Bash, or similar languages
- Demonstrated ability to produce clear, structured technical design documentation
- Excellent written and verbal communication skills
- Bachelor’s degree in computer science or equivalent professional experience
Preferred Qualifications
- OpenStack, Ubuntu KVM administration
- BareMetal as a Service (PXE, Redfish)
- Kubernetes on BareMetal
- CIS/NIST security and infrastructure lifecycle management
- ITIL Foundation/advanced certifications in support of ITSM standard methodology
- Background in telco, edge cloud, or large enterprise environments
- Ubuntu Certifications, CNCF Certified Kubernetes Administrator (CKA), Certified Kubernetes Security Specialist (CKS)
- Master’s degree in computer science, IT, Engineering, or a related field preferred; equivalent experience and relevant industry certifications will also be considered
Skills and Attributes
- Analytical Thinking & Problem Solving: Demonstrated ability to translate complex, cross-domain requirements into scalable and resilient cloud infrastructure and automation solutions
- Collaboration & Teamwork: Strong interpersonal and communication skills with a proven track record of effective collaboration across multidisciplinary teams, including developers, operations, security, and product stakeholders
- Mentorship & Leadership: Passionate about knowledge-sharing and mentorship, with experience guiding junior engineers and fostering a team culture of continuous learning, innovation, and technical excellence in cloud engineering and DevOps practices
Schedule
This role is fully remote for candidates who reside outside the 30-mile radius of one of our offices. For candidates who reside within a 30-mile radius of one of our offices, this role is Hybrid and would require 3 days a week (T, W, TH) in office.
Pay
The US base salary range for this role is $82,300—$228,800 USD. Actual compensation packages are based on several factors that are unique to each candidate including, but not limited to: skill set, depth of experience, certifications, and specific work location. The range displayed reflects the minimum and maximum target for new hire salaries for the job across the United States. Additionally, the total compensation package for this position may also include an annual performance bonus, stock, and/or other applicable incentive compensation plans.
Benefits
- Health, dental, and vision coverage, beginning on the first day of employment (Five9 covers 100% of the employee portion and shares a high portion of the dependent cost)
- Short & Long-Term Disability, Basic Life Insurance, and a 401k savings plan with employer matching
- Access to an innovative mental health support platform that offers personalized care and resources in areas such as therapy, coaching, and self-guided mindfulness exercises for all covered employees and their covered dependents
- Generous employee stock purchase plan
- Paid Time Off, company-paid holidays, paid volunteer hours, and 12 weeks paid parental leave