Senior Platform Engineer, Network Infrastructure - DGX Cloud
About the role
The team owns the architecture and lifecycle of the Kubernetes platform, including cluster provisioning and upgrades, GitOps delivery, observability, capacity, and service enablement. The role involves owning the lifecycle management for GNI Kubernetes environments, developing production-quality software and automation for cluster provisioning, validation, upgrades, remediation, and safe multi-cluster delivery through GitOps. Production support for network services hosted on the platform is also provided, along with diagnosing complex Kubernetes platform and hosted-service failures.
Responsibilities
- Design, build, and operate the Kubernetes platform that powers GNI network automation, telemetry, and operations across data center, colocation, and cloud environments.
- Own the lifecycle management for GNI Kubernetes environments, including cluster onboarding, upgrades, capacity, availability, and recovery.
- Develop production-quality software and automation for cluster provisioning, validation, upgrades, remediation, and safe multi-cluster delivery through GitOps.
- Provide production support for network services hosted on the platform, working with Network Automation and service teams that retain ownership of application architecture, code, and features.
- Diagnose complex Kubernetes platform and hosted-service failures involving control-plane health, cluster networking, storage, scheduling, workload placement, and multi-cluster dependencies. Drive issues from initial signal through verified resolution.
- Define production-readiness and observability standards for the platform and hosted network services, including health signals, capacity, alerts, runbooks, and recovery.
- Participate in CFR’s production on-call rotation, including scheduled after-hours and weekend coverage. Lead incident response and recovery, then drive corrective actions to completion.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.
- 8+ years of experience building or operating production Kubernetes platforms, network infrastructure, or distributed systems.
- Deep experience with Kubernetes at scale, including cluster lifecycle, upgrades, networking, storage, and recovery.
- Proficiency in at least one general-purpose programming language, such as Go or Python.
- Experience with GitOps, infrastructure as code, CI/CD, and automated production delivery.
- Experience deploying and supporting network automation or telemetry services on Kubernetes.
- Experience with production on-call, incident response, root-cause analysis, and driving corrective actions to completion.
Qualifications
- Strong knowledge of IP routing, data center fabrics, and cloud networking is a plus.
- Hands-on experience with Cluster API (CAPI) and Metal3 for bare-metal provisioning, cluster lifecycle, machine remediation, and upgrades.
- Experience building Kubernetes controllers or operators in Go using custom resources and reconciliation patterns.
- Experience designing or operating network automation and telemetry services on Kubernetes at global scale.
- Contributions to Cluster API, Metal3, or other open-source Kubernetes infrastructure projects.
Skills
- Deep understanding of Kubernetes at scale.
- Experience with GitOps, CI/CD, and automated production delivery.
- Hands-on experience with Kubernetes controllers or operators in Go.
- Experience with Cluster API (CAPI) and Metal3.
- Experience with network automation and telemetry services on Kubernetes.
Benefits
- Competitive base salary ranging from $176,000 to $333,500 based on experience and location.
- Equity and comprehensive benefits package.
Pay
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is $176,000 - $276,000 for Level 4, and $208,000 - $333,500 for Level 5.
Schedule
This is a full-time position with a production responsibility and end-to-end ownership.