Distinguished Engineer, End-to-End Scaling Performance Architecture
NVIDIA · Redmond, WA · Today
EngineeringFull-time
About the role
Join NVIDIA's architecture organization to help define how future accelerated computing systems scale from a single processor to multi-die, multi-GPU, and multi-node platforms. Set long-term performance strategy across applications, systems, and architecture, including DRAM, NVLink, and chip-to-chip (C2C) interconnects. Focus on architectural direction and application outcomes.
Key responsibilities
- Define the multi-generation strategy for application scaling across DRAM, NVLink, C2C, compute, and the supporting software stack.
- Translate the behavior of important AI, HPC, and accelerated computing applications into architectural requirements, performance targets, and investment priorities.
- Build a clear view of how bottlenecks shift as workloads scale across dies, GPUs, nodes, model sizes, data sets, and communication patterns.
- Evaluate system-level trade-offs across bandwidth, latency, capacity, topology, coherence, power, area, cost, programmability, and resiliency.
- Establish common workload scenarios, scaling metrics, models, and decision frameworks so architecture teams can compare proposals against application outcomes.
- Identify architectural discontinuities and emerging technology opportunities early enough to shape product and technology decisions.
- Partner with DRAM, NVLink, C2C, GPU, CPU, system, and software architects to align around shared performance limits and high-value opportunities.
- Work with application, framework, compiler, runtime, modeling, and post-silicon teams to connect measured behavior with future architecture choices.
- Provide clear recommendations to senior technical and business leaders, including assumptions, sensitivities, risks, and expected impact.
- Mentor system performance architects, strengthen technical communities across teams, and generate sustained intellectual property.
Qualifications
- MSEE, MSCE, PhD, or equivalent experience in Electrical Engineering, Computer Engineering, Computer Science, or a related field.
- 18+ years of relevant industry or academic experience, including experience setting architecture direction for complex, high-performance systems.
- Deep understanding of system performance and scaling, including interactions among DRAM behavior, high-bandwidth fabrics such as NVLink, and C2C communication.
- Strong application-level intuition, including the ability to connect workload algorithms, parallelism, communication, locality, and data movement to architecture choices and measurable outcomes.
- Experience with workload characterization, analytical or simulation-based performance modeling, bottleneck analysis, and architecture trade-off evaluation.
- Record of identifying cross-domain opportunities that may not be visible when teams optimize individual components separately.
- Demonstrated ability to create and advance a multi-generation technical strategy through influence across silicon, systems, software, and application teams.
- Clear communication and sound judgment in ambiguous technical areas, with the ability to explain complex system trade-offs to specialists and executive leaders.
- Experience mentoring senior engineers into broader architecture leadership roles and building strong technical communities.
Preferred qualifications
- Shaped product or technology roadmaps around application-level scaling needs.
- Helped architecture teams build a quantitative understanding of bottlenecks across memory, interconnect, compute, and software.
- Identified high-impact trade-offs early enough to guide product, architecture, or technology investment.
- Delivered measurable end-to-end improvements in performance, efficiency, or scaling for priority applications.
Pay
Base salary range: 320,000 USD - 488,750 USD. Also eligible for equity and benefits.
Schedule
Applications accepted at least until August 1, 2026. This posting is for an existing vacancy.