Jobs · Information Technology · California

Senior AI Systems and Algorithms Engineer

NVIDIA AI · Santa Clara County, CA · 3 wk ago
Information TechnologyFull-time

NVIDIA is seeking a Senior GenAI Algorithms Engineer to advance the state of the art in foundation model development, training, and deployment. This role spans the entire GenAI lifecycle from large-scale data preparation to training, post-training, inference optimization, and framework development.

Responsibilities

  • Data Curation & Readiness: Design scalable systems for preparing high-quality multimodal datasets for frontier foundation model training.
  • Training Efficiency: Develop algorithms and systems that improve the scalability, efficiency, and cost of large-scale pre-training and post-training.
  • Inference Efficiency: Advance techniques that improve inference performance, reduce deployment cost, and enable efficient serving across cloud and edge platforms.
  • Open-Source AI Infrastructure: Develop reusable infrastructure and contribute brand new model support to NVIDIA's open-source GenAI training platform.

Requirements

  • MS or Ph.D in Computer Science, AI, Applied Mathematics, or a related field (or equivalent experience).
  • 5+ years of relevant industry experience.
  • Strong foundation in machine learning, deep learning, and optimization.
  • Excellent software engineering skills, including Python and PyTorch.
  • Experience building high-performance software for large-scale AI systems.
  • Strong analytical, debugging, and performance optimization skills.
  • Excellent communication and collaboration skills.

Skills

Experience in some of the following areas is highly desirable:

  • Large-Scale Training: Distributed training at scale, including Megatron-LM, Megatron Bridge, FSDP, TP/PP/CP/DP, heterogeneous or per-module parallelism, optimizer research, and efficient sparse or long-context attention.
  • LLM/VLM Post-Training: Supervised fine-tuning (SFT), reinforcement learning for LLMs (e.g., PPO, GRPO, asynchronous RL), and large-scale RL frameworks such as NeMo-RL.
  • Inference Efficiency: Model compression techniques including quantization (FP8, NVFP4, INT4), pruning, knowledge distillation, neural architecture search, and diffusion or non-autoregressive language models.
  • Open-Source AI Infrastructure: Contributing to open-source AI frameworks such as Megatron-LM, Megatron Bridge, NeMo-RL, or Hugging Face Transformers along with experience in GPU performance optimization, distributed systems, latency/throughput analysis, and profiling of large-scale AI workloads.

Pay

The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4. You will also be eligible for equity and benefits.

Similar jobs

Senior AI Engineer

Carzma AISan Francisco, CA· 1 mo ago
Information Technologyapply on carzma.ai

Senior AI Engineer

Tata Consultancy ServicesOakland, CA· 1 mo ago
Engineeringapply on ibegin.tcsapps.com