Jobs · Quality Assurance

DevOps (Cloud/AI)

Bayside Solutions · Cupertino, CA · 1 mo ago
RemoteRemoteQuality Assurance$60–$70/hrContract

Duties and Responsibilities

  • Design, build, automate, and support scalable Kubernetes-based platforms and services
  • Operate and troubleshoot production environments running at scale
  • Develop automation and tooling to improve operational efficiency and reliability
  • Maintain platform health, performance, and availability using observability tooling
  • Identify and resolve infrastructure, application, and networking issues across distributed systems
  • Collaborate with engineering teams to improve deployment, reliability, and scalability practices
  • Support operational support, incident response, and root cause analysis
  • Optimize CI/CD workflows and deployment automation
  • Drive operational excellence through documentation, automation, and process improvements
  • Take ownership of projects and independently drive deliverables to completion

Requirements and Qualifications

  • Strong hands-on experience with Kubernetes platforms such as EKS, GKE, and AKS or similar
  • Experience running and supporting applications on Kubernetes at scale
  • Strong understanding of containerized infrastructure and distributed systems
  • Experience with monitoring and observability tools, preferably Grafana or Prometheus
  • Experience with CI/CD pipelines and deployment automation
  • Experience with Splunk logging, log analysis, and troubleshooting
  • Strong scripting and automation experience using Python and/or Golang
  • Experience troubleshooting production systems under pressure
  • Effective communication and collaboration skills
  • Self-starter mentality with strong ownership and accountability

Preferred Qualifications

  • Experience operating Ray clusters/services
  • Networking and troubleshooting experience
  • Experience with cloud infrastructure and platform services
  • Experience with Infrastructure as Code and automation frameworks
  • Experience supporting high-scale production systems
  • Familiarity with SRE principles and operational best practices

Desired Skills and Experience

  • Kubernetes, DevOps, Site Reliability Engineering, Cloud Infrastructure, AI Infrastructure, Platform Engineering
  • EKS, GKE, AKS, Containerized Infrastructure, Distributed Systems, Production Operations, Infrastructure Automation, CI/CD Pipelines, Deployment Automation, Infrastructure as Code
  • Python, Golang, Grafana, Prometheus, Splunk, Log Analysis, Monitoring and Observability, Ray Clusters, Network Troubleshooting, Incident Response, Root Cause Analysis, Performance Monitoring, High Availability, Scalability, Reliability Engineering, Operational Excellence, Technical Documentation, Cross-Functional Collaboration, Project Ownership

Similar jobs

DevOps/AI Engineer

MITREMcLean, VA· 1 mo ago
Engineering$86k–$108k/yrapply on careers.mitre.org

Cloud / AI Developer

General Dynamics Information TechnologyUnited States· 3 wk ago
RemoteEngineering$128k–$173k/yrapply on gdit.wd5.myworkdayjobs.com