Senior Software Engineer, Networking DGX Cloud
NVIDIA · United States · 1 wk ago
RemoteRemoteEngineeringFull-time
About the role
We are looking for an experienced software engineer with infrastructure experience to become a senior member of the Cloud Foundations Automation - Development Team. We build and manage the automation ecosystem supporting NVIDIA's GPU Cloud and NVIDIA SuperPod deployments.
Responsibilities
- Develop software to enable efficient network design, deployment, and day 2 management.
- Build product-focused software solutions used by internal and external customers.
- Transform workflows and organization into a centrally orchestrated configuration management framework, operating at scale across geographies.
- Own and drive integrations with various service APIs such as Cloud Service Providers to automate creation of environments and auto-populate data sources.
- Build on open-source software, designing and implementing data structures and UI interfaces to automate processes from equipment purchase to device config generation to deployment to operations.
- Streamline deployment mechanisms and lifecycle operations.
- Develop modern service architectures around streaming data and event pipelines.
- Work with infrastructure domain experts on true zero-touch deployment solutions and utilize best-of-breed high-performance computing management solutions.
- Be a proactive problem solver, identifying new opportunities to improve services and customer experience.
- Communicate readily with peers across the organization, building relationships and collaborating enthusiastically.
Requirements
- BS or equivalent experience with 12+ years of relevant industry experience.
- Background in networking and network automation (Datacenter / PoPs / Routing).
- Strong proficiency in Python and Go web frameworks.
- Experience building and shipping production-quality software products.
- Experience with DCIM tools like Netbox or Nautobot and expertise in designing and implementing network configuration management systems.
- Familiarity with containerization using Kubernetes (on-premises, EKS) and streaming telemetry protocols like gRPC/GNMI.
Skills
- Architected, built, and deployed a network for a large scale (1000s of machines) used by hundreds of engineers.
- Hands-on experience building tooling and automation for provisioning, monitoring, and managing network infrastructure.
- Understanding of networking technologies such as VRFs, VxLAN, EVPN, BGP/OSPF/ISIS, and datacenter design.
Pay
The base salary range is 200,000 USD - 322,000 USD for Level 5, and 248,000 USD - 391,000 USD for Level 6. You will also be eligible for equity and benefits.
JR2022482