Senior Solutions Architect, Generative AI Deployment and AIOps
About the role
NVIDIA is seeking outstanding AI Solutions Architects to assist and support customers building solutions with our newest AI technology. At NVIDIA, solutions architects work across different teams and help customers with the latest Accelerated Computing and Deep Learning software and hardware platforms. You will become a trusted technical advisor, working on exciting projects and proof-of-concepts focused on inference for Generative AI and Large Language Models (LLMs). You will also collaborate with internal teams on performance analysis and modeling of inference software. This role offers the opportunity to work in an interdisciplinary team at the forefront of technological advancement.
Responsibilities
- Partner with other solution architects, engineering, product, and business teams to understand strategies and technical needs, helping define high-value solutions
- Dynamically engage with developers, scientific researchers, and data scientists, gaining experience across a range of technical areas
- Strategically partner with lighthouse customers and industry-specific solution partners targeting our computing platform
- Work closely with customers to help them adopt and build creative solutions using NVIDIA technology and MLOps solutions
- Analyze performance and power efficiency of AI inference workloads on Kubernetes
- Travel to conferences and customers may be required (20%)
Requirements
- BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, other Engineering or related fields (or equivalent experience)
- 8+ years of hands-on experience with Deep Learning frameworks such as PyTorch and TensorFlow
- Strong fundamentals in programming, optimizations, and software design, especially in Python
- Proficiency in problem-solving and debugging skills in GPU orchestration and Multi-Instance GPU (MIG) management within Kubernetes environments
- Experience with containerization and orchestration technologies, monitoring, and observability solutions for AI deployments
- Excellent knowledge of the theory and practice of LLM and DL inference
- Excellent presentation, communication, and collaboration skills
Skills
- Prior experience with DL training at scale, deploying or optimizing DL inference in production
- Experience with NVIDIA GPUs and software libraries such as NVIDIA NIM, Dynamo, TensorRT, TensorRT-LLM
- Excellent C/C++ programming skills, including debugging, profiling, code optimization, performance analysis, and test design
- Familiarity with parallel programming and distributed computing platforms
Pay
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD. You will also be eligible for equity and benefits.