Jobs · Engineering · Washington

Senior Software Engineer, AI Resiliency

NVIDIA · Redmond, WA · 1 mo ago
EngineeringFull-time

About the role

We are now looking for a Senior Software Engineer for AI Resiliency at NVIDIA. Our team is focused on developing AI software resiliency for the most powerful AI supercomputers in the world.

Responsibilities

  • Develop AI Software Resiliency Features: Implement and optimize software features that improve AI system reliability at a massive scale, such as fast checkpoint-recovery, error detection, error isolation, and straggler/hang detection.
  • Hands-On Coding & Optimization: Contribute to large-scale distributed systems with high-quality, production-level C++ and Python code. Enhance performance for AI workloads running on thousands of GPUs.
  • Fault Tolerance & Debugging: Work on AI system error handling, implementing techniques to detect silent data corruption (SDC) and other failure scenarios. Assist in developing monitoring tools for proactive failure mitigation.
  • Collaborate Across Teams: Work closely with senior engineers, AI researchers, and hardware/software teams to integrate resiliency features into AI frameworks like PyTorch and JAX/XLA.
  • Testing & Automation: Develop and implement tests to ensure robustness, scalability, and efficiency of resiliency mechanisms. Contribute to CI/CD pipelines to automate validation of AI workloads.
  • Support Production Deployments: Assist in debugging and performance tuning large-scale AI workloads in cloud and HPC environments, ensuring seamless operation of AI training and inference workloads.

Requirements

  • Has achieved a Bachelor’s, Master’s or PhD in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • Proficiency in C++ and Python, with experience in writing efficient, high-performance code.
  • 6+ years of relevant experience.
  • Strong understanding of distributed systems concepts, parallel programming, and fault tolerance in large-scale computing environments.
  • Familiarity with AI frameworks such as PyTorch, JAX/XLA, TensorFlow, or similar.
  • Experience with debugging and profiling tools (e.g., gdb, perf, valgrind, NVIDIA Nsight).
  • Excellent problem-solving skills and ability to work in a fast-paced, highly collaborative environment.

Qualifications

  • Hands-on experience in training models or working with model training teams.
  • Hands-on experience with CUDA, NCCL, or MPI for GPU-accelerated computing, especially at extreme-scale.
  • Knowledge of checkpointing strategies, error mitigation, or fault-tolerant computing in AI training.
  • Experience working with large-scale AI clusters, HPC environments, or cloud-based AI workloads.
  • Strong systems programming skills and experience with low-level performance tuning.

Skills

  • Experience with AI tools and technologies.
  • Ability to work in a dynamic, fast-paced environment.
  • Strong communication and collaboration skills.

Benefits

Base salary range: $184,000 - $287,500. Eligible for equity and benefits.

Pay

Base salary will be determined based on your location, experience, and the pay of employees in similar positions.

Schedule

Full-time position.

Similar jobs