Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
What You'll Own
This is the Kubernetes-native deployment platform paired with its automation system and the backend services and APIs that support it. We deliver zero-downtime global rollouts with automatic drift detection. Our reliable staging environments replicate production accurately. We provide cloud automation for bare-metal and multi-cloud environments. If a GFN service ships, runs, or self-heals, you build the infrastructure behind it.
What You'll Be Doing
- Build and develop backend microservices and REST/gRPC/MCP APIs that power the deployment platform, zone reservation/lease system, and developer self-service tooling
- Extend the platform with dynamic delivery, automatic rollback, drift detection, and automated zone bootstrapping
- Build and stabilize prod-like staging environments/zones so teams catch regressions early
- Develop and sustain GitOps pipelines (Flux CD, Argo CD) across on-premises and Nvidia GFN Cloud/AWS/Azure/GCP
- Develop Kubernetes CRDs and operators in Go for scheduling, auto-scaling, and compliance across data centers
- Build backend integrations and control-plane services connecting CI/CD, observability, and automation systems into a unified platform experience
- Automate dedicated hardware and multiple cloud platform configurations using Terraform, Ansible, and Vault
- Implement monitoring solutions including Prometheus, Grafana, Datadog, and ELK, paired with SLO/alerting for early detection
- Integrate automation tools — runbooks, StackStorm bots, anomaly-triggered remediation, and Slack self-service release bots
What We Need To See
- Bachelor or higher degree in computer science, engineering, or equivalent experience
- 10+ years of experience in Cloud Infrastructure and DevOps, with deep expertise in Kubernetes, GitOps (or equivalent), and production-grade cloud-native CI/CD pipelines
- Expert Kubernetes: CRDs, operators, multi-cluster management, and security hardening (CIS, PCI/SOC 2)
- Proficiency in Flux CD or Argo CD, GitLab CI, and Jenkins; Go or Python for control plane development
- Experience handling Vault, Terraform, Ansible, and Helm in both on-premises and cloud environments
- Experience in developing and scaling RESTful, gRPC, MCP APIs and backend services
- Experience running hybrid multi-cloud and bare-metal at production scale, and owning a platform roadmap end-to-end
Ways To Stand Out From The Crowd
- Backstage or similar internal developer portals for self-service tooling
- Comprehensive expertise covering backend services as well as platform and infrastructure engineering
- Daily use of AI-assisted tools (Claude Code, GitHub Copilot, Cursor)
- Hybrid infrastructure spanning on-premises GPU clusters and public cloud
Pay & Benefits
Competitive salaries and a generous benefits package are offered. Applications for this job will be accepted at least until July 30, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.