Senior Systems Software Engineer, Compute Stack Acceleration
About the Role
NVIDIA is growing a senior engineering team focused on making our compute software stack first-class on NVIDIA CPU platforms. The team turns modern toolchains, build and code-health practices, performance-analysis workflows, and optimization techniques into repeatable improvements across real software components. We are looking for an experienced systems software engineer who can lead cross-stack engineering efforts from an ambiguous adoption problem to a measurable outcome. You will work closely with compiler, platform, performance, library, and application teams to validate new capabilities, resolve integration blockers, produce credible before-and-after evidence, and turn successful approaches into reusable engineering practices. Depending on your background and project needs, your initial work may emphasize toolchain, build, and code-health adoption or profiling, optimization, and evidence-routing workflows. The role is intentionally broad enough to evolve as the team identifies the highest-leverage opportunities across the stack.
What You'll Be Doing
- Lead toolchain, build, code-health, and performance-workflow adoption projects across large software components.
- Evaluate and deploy supported GCC and LLVM/Clang toolchains, compiler options, linkers, sysroots, and cross-compilation configurations.
- Integrate modern toolchains and workflows into complex build systems and continuous-integration environments.
- Establish useful Clang diagnostic builds and targeted sanitizer coverage in partnership with component owners.
- Evaluate techniques such as link-time optimization, profile-guided optimization, AutoFDO, and BOLT on representative software.
- Use profiling, PMU data, flamegraphs, and binary/source analysis to identify actionable performance and code-quality findings.
- Build automation, wrappers, validation scripts, dashboards, and migration helpers where they improve adoption and repeatability.
- Measure changes in runtime, code size, build time, launch latency, throughput, quality, or engineering velocity.
- Drive complex blockers to the appropriate compiler, runtime, library, infrastructure, or component owner.
- Document validated approaches as reusable playbooks for other engineering teams.
- Communicate technical results and tradeoffs clearly to engineers and leadership.
Requirements
- BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent experience.
- 8+ years of relevant systems software engineering experience.
- Strong C and C++ development, debugging, and code-review skills.
- Strong Linux systems knowledge and hands-on experience with complex native software stacks.
- Practical experience with GCC or LLVM/Clang, linkers, compiler options, and build systems.
- Experience working in large, multi-component codebases and CI environments.
- Experience with performance profiling, root-cause analysis, and before-and-after validation.
- Ability to lead projects with substantial technical and organizational ambiguity.
- Excellent written and verbal communication and a record of effective cross-team collaboration.
Preferred Qualifications
- Experience with Arm64 systems, CPU architecture, vectorization, or SVE.
- Experience with LTO, PGO, AutoFDO, BOLT, binary optimization, or code-layout analysis.
- Hands-on experience with Clang diagnostics, AddressSanitizer, or other code-health workflows.
- Familiarity with PMU analysis, perf, flamegraphs, BRBE, SPE, ETM, or similar profiling technologies.
- Experience with cross-compilation, sysroots, large monorepositories, Perforce-scale development, or distributed build systems.
- Experience turning a successful migration or optimization into a maintained workflow used by multiple teams.
- Python or other scripting experience for engineering automation and data analysis.
Pay
Base salary range: 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits.