Jobs · Engineering · Texas

Senior Kubernetes Platform Engineer - AI/ML Infrastructure

Cisco · Allen, TX · 2 wk ago
Engineering$139k–$204k/yrFull-time

Your Impact / Core Responsibilities

  • Architect, build, and operate large-scale on-prem Kubernetes platforms (OpenShift/Anthos), including control plane and etcd lifecycle management
  • Define and evolve scalable, multi-tenant platform architecture supporting AI/ML and GPU-based workloads
  • Enable and optimize ML workloads including training, inference, and LLM deployment pipelines on Kubernetes
  • Build platform extensions using Kubernetes controllers, operators, CRDs, and Golang-based services
  • Implement Infrastructure as Code and automation to improve scalability, consistency, and operational efficiency
  • Drive AIOps capabilities using telemetry, automation, anomaly detection, and self-healing systems for platform reliability
  • Improve observability (metrics, logs, traces) and optimize resource utilization, scheduling, and cluster performance
  • Partner with ML engineers and data scientists to operationalize ML workflows and ensure platform readiness for AI workloads
  • Participate in on-call rotations, owning incident response, reliability, and continuous operational improvement
  • Mentor engineers and contribute to defining platform engineering standards and best practices

Minimum Qualifications

  • 8+ years of software engineering experience
  • 4+ years of hands-on Kubernetes production experience with control plane ownership
  • Strong experience operating on-prem or self-managed Kubernetes environments
  • Deep expertise in etcd management (backup, restore, recovery, upgrades)
  • Strong proficiency in Go with experience building Kubernetes controllers, operators, CRDs, and webhooks
  • Deep understanding of Kubernetes internals (API server, scheduler, controller loops, reconciliation)
  • Experience supporting AI/ML or GPU-based workloads on Kubernetes platforms
  • Proven experience operating and debugging large-scale distributed systems
  • Experience participating in on-call rotations and production incident management

Preferred Qualifications

  • Experience with bare-metal or enterprise on-prem infrastructure at scale
  • Exposure to AI/ML platforms and tooling (e.g., Kubeflow, MLflow, distributed training systems)
  • Experience building internal developer platforms or platform-as-a-service (PaaS) systems
  • Familiarity with AIOps concepts such as automated remediation and predictive operations
  • Experience applying data-driven or ML-based techniques for system reliability or capacity planning
  • Contributions to Kubernetes, CNCF, or other open-source ecosystems

Similar jobs