Manager, DevOps ()
CareersElite.com · San Jose, CA · 1 wk ago
Management$66k–$99k/yrFull-time
San Jose, California, United States
About the role
Architect, build, and continuously evolve a truly declarative, self-healing, cloud-native, Kubernetes-native, and highly observable enterprise-grade platform. This strategic leadership role defines and drives the fundamental architecture of our technology stack, including:
- Go-based microservices architecture, ensuring high throughput and resilience
- Next-generation CI/CD pipelines utilizing GitHub Actions and Azure DevOps, engineered for speed, reliability, governance, and release integrity
- Infrastructure-as-Code (IaC) automation using Terraform and Ansible to deliver consistent, secure, and repeatable environments at scale
- World-class Observability strategy, integrating metrics, logging, and tracing, and intelligent alerting to enable proactive operations and optimization
- Governance of MongoDB data persistence and distributed systems, ensuring performance, durability, and operational efficiency
Responsibilities
CI/CD Strategy & Governance (The Quality Pipeline)
- Responsibility & Authority (R&A) for the end-to-end architecture, standardization, and all YAML-based delivery pipelines (build, test, security scanning, artifact management)
- Establish and enforce mandatory, non-negotiable quality gates (unit, integration, security, performance) at every stage of the pipeline to ensure "shift-left" quality
- Secure and harden the CI/CD ecosystem against tampering and supply chain risks through strong access controls, artifact integrity, and policy enforcement
Deployment Excellence & Safety (Production Integrity)
- Define, standardize, and govern safe, high-integrity deployment strategies (e.g., Canary Releases, Rolling Updates, Blue/Green, utilization of Feature Flags)
- Implement robust, fast, and fully automated rollback mechanisms, ensuring absolute production safety and minimizing Mean Time To Recovery (MTTR)
- Continuously reduce deployment risk and complexity through automation, standardization, and release discipline
Production Reliability Engineering (SRE & Resilience)
- Establish, evangelize, and enforce enterprise-grade standards for Stability, Performance, Scalability, Observability, and Disaster Recovery (DR)
- Develop, monitor, and report on key Service Level Objectives (SLOs) and Service Level Agreements (SLAs)
- Lead and mature the incident response lifecycle, focusing on root cause analysis and preventative, long-term remediation
AI-Accelerated DevOps (Responsible Innovation)
- Champion the responsible and ethical integration of generative AI tools to accelerate repeatable engineering tasks such as IaC generation, test case development, and documentation creation
- Implement prompt engineering practices to maximize AI effectiveness and reliability
- Establish and enforce rigorous verification and validation processes for all AI-generated artifacts
Platform Architecture & Optimization (Scalability & Performance)
- Oversee the continuous evolution of the scalable, resilient architecture for the foundational Kubernetes control plane
- Drive the design of resilient Go microservices and REST APIs, emphasizing security and performance
- Provide expertise in optimizing MongoDB performance, indexing strategies, and distributed systems design to ensure 24/7 availability and data integrity
Requirements
- BS degree in Computer Science, Computer Engineering or related field
- 10 years of experience as an Engineering Manager, Software Engineer, or related role
Skills
- 10 years of experience in software engineering, DevOps, or Site Reliability Engineering (SRE)
- 5 years of programming experience in Python, Go, Java, or a comparable enterprise-grade language
- 5 years of experience operating, scaling, and troubleshooting Kubernetes within large, complex production environments
- 5 years of experience with MongoDB schema design, query optimization, and performance tuning for high volume, mission-critical environments
- 5 years of experience designing, governing, and securing modern CI/CD pipelines (e.g., Jenkins, GitHub Actions, Azure DevOps) with strong emphasis on automation and compliance controls
- 5 years of experience with Infrastructure-as-Code (Terraform and/or Ansible) and with one or more major cloud platforms (AWS, Azure, or GCP)
- 3 years of direct technical leadership experience, leading teams supporting high-scale distributed systems