Principal Software Engineer – Infrastructure
NVIDIA · Redmond, WA · Today
EngineeringFull-time
About the role
NVIDIA is seeking a Principal Software Engineer to lead the development of foundational software systems that manage infrastructure consistently across various environments. This role involves defining the multi-year technical vision and architecture for IT infrastructure automation, configuration management, orchestration, and self-service platforms.
Responsibilities
- Define the multi-year technical vision and architecture for IT infrastructure automation, configuration management, orchestration, and self-service platforms.
- Establish the technical direction for infrastructure automation and configuration management across technologies such as Ansible Automation Platform, AWX, Salt or equivalent platforms.
- Establish the strategy for applying AI to configuration management, including configuration intelligence, drift and compliance analysis, change-risk identification, root-cause assistance, intelligent recommendations, and guarded automated remediation.
- Develop prototypes and production software, review critical code and designs, and resolve the most challenging technical and scalability problems.
- Build secure and scalable integrations across infrastructure platforms, cloud services, CMDB, secrets management, observability, and enterprise data systems.
- Ensure platforms meet enterprise requirements for availability, scalability, performance, security, disaster recovery, observability, and operational support.
- Lead organizations in implementation and adoption within large-scale environments.
- Bring deep cross-domain technical expertise to identify gaps and challenge current methods. Establish architectures and engineering standards. Lead organizations in implementation and adoption within large-scale environments.
- Partner with engineering and executive leadership to translate critical business challenges into technical strategy, prioritized roadmaps, and measurable business outcomes.
Requirements
- 15+ years of progressive software engineering experience, with a sustained record of delivering complex, business-critical platforms and distributed systems.
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a similar domain, or a Master’s degree or equivalent experience.
- Deep knowledge of enterprise configuration management or infrastructure automation using technologies such as Ansible Automation Platform, AWX, Salt or equivalent platforms.
- Strong hands-on experience in Linux, Kubernetes, containers, cloud and hybrid infrastructure, CI/CD and Infrastructure as Code.
- Proven experience defining, owning, and evolving the architecture of large-scale infrastructure platforms operating across multiple teams, data centers, or cloud environments.
- Demonstrated ability to identify organization-wide business and technical challenges, establish a clear strategy, and drive implementation across teams.
- Outstanding communication and technical leadership skills, with a proven ability to build consensus, influence senior leaders, mentor experienced engineers, and lead through ambiguity.
Qualifications
- Experience in architecting, deploying and managing Ansible Automation Platform at enterprise scale.
- Experience developing platforms that manage large global infrastructure fleets across data centers, public clouds, compute, storage, and networking environments.
- A history of leading the development and enterprise-wide adoption of an AI-enabled configuration management, infrastructure automation, or autonomous operations platform.
- Experience using AI to solve infrastructure challenges such as configuration drift, compliance, change-risk analysis, incident diagnosis, capacity management, predictive operations, or automated remediation.
Skills
- Experience with Go, Python, Java, or a comparable systems programming language, including APIs, concurrency, distributed systems, testing, debugging, and performance engineering.
- Strong hands-on experience in Linux, Kubernetes, containers, cloud and hybrid infrastructure, CI/CD and Infrastructure as Code.
- Proven experience defining, owning, and evolving the architecture of large-scale infrastructure platforms operating across multiple teams, data centers, or cloud environments.
- Demonstrated ability to identify organization-wide business and technical challenges, establish a clear strategy, and drive implementation across teams.
- Outstanding communication and technical leadership skills, with a proven ability to build consensus, influence senior leaders, mentor experienced engineers, and lead through ambiguity.
Benefits
- Comprehensive benefits package including health insurance, retirement plans, and paid time off.
- Equity and other compensation packages.