Principal Engineer, Applied Research - Accelerator Programming Model and Compiler
NVIDIA · Santa Clara, CA · 5 days ago
EngineeringFull-time
What You Will Be Doing
- Define the next-generation PVA programming model for developers and AI coding agents, making it easier to build optimized algorithms for PVA.
- Enable AI agent-driven PVA development by abstracting hardware-specific details into agent-accessible, declarative interfaces, supported by improvements to compile time, emulator speed, and diagnostic tooling.
- Provide technical leadership for the architecture, and feature set of the PVA SDK, runtime APIs and programming model.
- Define benchmarks and evaluations for agent-generated PVA code and use them to improve the compiler optimizations and diagnostics.
- Develop and actively improve the LLVM-based VPU compiler backend targeting VLIW/SIMD architecture.
- Research and define efficient integration models for PVA workloads in CUDA-based heterogeneous pipelines, including execution, memory movement and synchronization.
- Drive deep technical integrations with internal and external customers to improve adoption and shape PVA runtime API for real-world workloads.
What We Need To See
- BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or related field, or equivalent experience.
- 15+ years of experience building high-performance, low-level systems software, accelerator software, embedded software, or compiler/toolchain infrastructure.
- Experience building higher-level programming abstractions, DSLs or compiler IRs that abstract hardware complexity while preserving performance.
- Experience with compiler, debugger, linker, or toolchain development, particularly using LLVM.
- Experience developing code with agents — Claude Code, Cursor, Open AI and inference SDKs.
- Experience integrating compilers and developer tools with AI coding agents or agentic development harnesses.
- Experience programming SIMD/VLIW processors.
- Excellent software development skills in C++, including low-level debugging and performance profiling.
- Experience with Linux or QNX development environments.
- Strong communication and social skills.
Ways To Stand Out From The Crowd
- Experience with CUDA, especially integrating accelerators into a CUDA-based heterogeneous compute pipeline.
- Background with MLIR, Halide, TVM, Triton, graph compilers, image-processing DSLs, or other declarative/compiler-based programming systems.
- Familiarity with ISO 26262 and IEC 61508 or equivalent quality/safety standards.
- Experience building agent harnesses and orchestrating long-running, stateful, multi-agent workflows.
Pay & Benefits
Competitive salaries and a generous benefits package. Base salary range is 272,000 USD - 431,250 USD. Eligible for equity and benefits.
Application Information
Applications for this job will be accepted at least until July 20, 2026. This posting is for an existing vacancy.