Jobs · Engineering · California

Senior Site Reliability Engineer Kubernetes Platform

Tata Consultancy Services · San Jose, CA · 4 wk ago
Engineering$64k–$130k/yrFull-time

Roles & Responsibilities

  • Design, build, and operate production-grade Kubernetes platforms in regulated environments
  • Improve system reliability through automation, thoughtful design, and continuous iteration
  • Define and drive Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to guide reliability decisions
  • Build and evolve CI/CD pipelines that are secure, scalable, and easy to use
  • Implement robust observability (metrics, logs, traces) to make systems understandable and actionable
  • Reduce operational toil by automating repetitive processes and improving workflows
  • Partner with security and compliance teams to meet FedRAMP High and IL5 requirements without sacrificing developer velocity
  • Support ATO processes, including documentation, controls implementation, and audit readiness
  • Participate in on-call rotations supporting customer requests and paging alerts
  • Participate in incident response, blameless postmortems, and continuous improvement efforts
  • Help shape a platform that engineers enjoy using

Qualifications

  • 10+ years of experience in SRE, DevOps, or infrastructure engineering
  • Strong experience running Kubernetes in production (EKS, AKS, GKE, or upstream)
  • Hands-on experience working in FedRAMP High and/or DoD IL5 environments
  • Solid understanding of cloud infrastructure, Linux systems, and networking fundamentals
  • Experience with Infrastructure as Code (Terraform preferred)
  • Familiarity with CI/CD systems (GitHub Actions, GitLab CI, Jenkins, ArgoCD)
  • Proficiency in scripting or programming (Python, Go)
  • Experience building or operating observability platforms (Prometheus, Grafana, OpenTelemetry, ELK)
  • Working knowledge of compliance frameworks (e.g., NIST 800-53, STIGs, RMF)

Nice to Have

  • Experience with service mesh technologies (Istio, Linkerd)
  • Familiarity with policy-as-code (OPA/Gatekeeper, Kyverno)
  • Experience with GitOps workflows
  • Exposure to multi-cluster or hybrid cloud architectures
  • Knowledge of FIPS-compliant systems or DoD Cloud SRG
  • Relevant certifications (CKA, CKS, cloud provider certs, Security+)

Similar jobs