DevOps (Cloud/AI)
Bayside Solutions · Cupertino, CA · 1 mo ago
RemoteRemoteQuality Assurance$60–$70/hrContract
Duties and Responsibilities
- Design, build, automate, and support scalable Kubernetes-based platforms and services
- Operate and troubleshoot production environments running at scale
- Develop automation and tooling to improve operational efficiency and reliability
- Maintain platform health, performance, and availability using observability tooling
- Identify and resolve infrastructure, application, and networking issues across distributed systems
- Collaborate with engineering teams to improve deployment, reliability, and scalability practices
- Support operational support, incident response, and root cause analysis
- Optimize CI/CD workflows and deployment automation
- Drive operational excellence through documentation, automation, and process improvements
- Take ownership of projects and independently drive deliverables to completion
Requirements and Qualifications
- Strong hands-on experience with Kubernetes platforms such as EKS, GKE, and AKS or similar
- Experience running and supporting applications on Kubernetes at scale
- Strong understanding of containerized infrastructure and distributed systems
- Experience with monitoring and observability tools, preferably Grafana or Prometheus
- Experience with CI/CD pipelines and deployment automation
- Experience with Splunk logging, log analysis, and troubleshooting
- Strong scripting and automation experience using Python and/or Golang
- Experience troubleshooting production systems under pressure
- Effective communication and collaboration skills
- Self-starter mentality with strong ownership and accountability
Preferred Qualifications
- Experience operating Ray clusters/services
- Networking and troubleshooting experience
- Experience with cloud infrastructure and platform services
- Experience with Infrastructure as Code and automation frameworks
- Experience supporting high-scale production systems
- Familiarity with SRE principles and operational best practices
Desired Skills and Experience
- Kubernetes, DevOps, Site Reliability Engineering, Cloud Infrastructure, AI Infrastructure, Platform Engineering
- EKS, GKE, AKS, Containerized Infrastructure, Distributed Systems, Production Operations, Infrastructure Automation, CI/CD Pipelines, Deployment Automation, Infrastructure as Code
- Python, Golang, Grafana, Prometheus, Splunk, Log Analysis, Monitoring and Observability, Ray Clusters, Network Troubleshooting, Incident Response, Root Cause Analysis, Performance Monitoring, High Availability, Scalability, Reliability Engineering, Operational Excellence, Technical Documentation, Cross-Functional Collaboration, Project Ownership