Platform Infrastructure SRE (Kubernetes / Cloud / IaC)
Bayside Solutions · Cupertino, CA · 2 wk ago
RemoteRemoteEngineering$60–$70/hrContract
W2 contract role based in Cupertino, CA (remote).
About the role
We are looking for a strong Senior Platform Infrastructure SRE / Software Engineer to support the productionization and operation of large-scale Kubernetes-based platform services across multiple cloud environments. This role is best suited for someone with deep Kubernetes and infrastructure experience who can take capabilities developed by a platform engineering team and make them repeatable, scalable, observable, reliable, and production-ready across many environments.
Responsibilities
- Productionize Kubernetes-based platform services developed by platform engineering teams.
- Deploy and operate platform infrastructure across multiple cloud and production environments.
- Build repeatable environment provisioning and deployment automation.
- Develop reusable infrastructure templates, blueprints, and deployment patterns.
- Provision Kubernetes clusters, cloud infrastructure, networking, and service dependencies.
- Configure environment-specific infrastructure, connectivity, and platform services.
- Deploy, validate, upgrade, and maintain platform services throughout their lifecycle.
- Build monitoring, alerting, dashboards, logging, health checks, and operational controls.
- Establish and validate production-readiness standards for new platform capabilities.
- Troubleshoot complex failures across applications, Kubernetes, infrastructure, networking, and distributed systems.
- Work closely with platform developers to understand application behavior and identify operational gaps before production rollout.
- Implement and maintain Infrastructure as Code using technologies such as Crossplane, Terraform, Pulumi, or CloudFormation.
- Build and maintain Helm-based Kubernetes packaging and deployment patterns.
- Design safe rollout, rollback, upgrade, recovery, and lifecycle-management processes.
- Validate platform capacity, availability, scalability, and reliability.
- Configure and troubleshoot DNS, load balancing, VPC networking, routing, service connectivity, security policies, and certificates.
- Improve operational automation and reduce manual environment-specific work.
- Document architecture, deployment patterns, operational procedures, and troubleshooting guidance.
- Participate in production support, incident response, and root-cause analysis as appropriate.
- Independently own technical work and drive complex problems through resolution with limited supervision.
Requirements
- Deep experience with Crossplane for infrastructure provisioning and platform automation.
- Experience with Alibaba Cloud and its Kubernetes, networking, and infrastructure services.
- Experience with AWS EKS and/or Google Cloud Platform.
- Experience designing or operating multi-cloud platforms.
- Experience with service mesh technologies and Kubernetes service networking.
- Experience with distributed data technologies such as Apache Spark, Apache Flink, or Trino.
- Experience implementing authentication, authorization, cloud security, and governance controls.
- Experience building and maintaining CI/CD pipelines for Kubernetes-based platforms.
- Experience designing highly available and resilient platform architectures.
- Experience automating provisioning and lifecycle management across a large number of environments.
- Strong production SRE, incident response, reliability engineering, and operational automation experience.
Skills
- Kubernetes
- Site Reliability Engineering (SRE)
- Platform Engineering
- Cloud Infrastructure
- Infrastructure as Code (IaC)
- Crossplane, Terraform, Pulumi, AWS CloudFormation
- Alibaba Cloud, AWS, Amazon EKS, Google Cloud Platform (GCP)
- Multi-Cloud Architecture
- Kubernetes Cluster Provisioning
- Helm
- CI/CD Pipelines
- Deployment Automation
- Infrastructure Automation
- Environment Provisioning
- Production Operations
- Production Readiness
- Monitoring, Alerting, Dashboards, Logging, Health Checks
- Observability
- Incident Response
- Root Cause Analysis
- Troubleshooting
- Distributed Systems
- Service Mesh
- Kubernetes Networking
- DNS, Load Balancing, VPC Networking, Routing, Service Connectivity
- Security Policies, Certificate Management
- Authentication, Authorization, Cloud Security, Cloud Governance
- High Availability, Resilient Architecture
- Scalability, Capacity Planning, Reliability Engineering
- Rollouts, Rollbacks, Platform Upgrades, Disaster Recovery, Lifecycle Management
- Apache Spark, Apache Flink, Trino
- Technical Documentation
- Production Support
- Software Engineering
Pay
$60 - $70 per hour.