Jobs · Engineering · California

Performance Engineer

Astera Labs · San Jose, CA · 1 mo ago
Engineering$135k–$170k/yrVolunteer

Key Responsibilities

  • Performance Characterization & Benchmarking
  • Build and maintain baseline performance benchmarks using industry-standard tools such as NVBandwidth and NCCL across a range of GPU configurations and switch topologies.
  • Quantify the impact of differentiated Astera Labs AI fabric features (e.g., Hypercast, In-Network Computing) against baselines using both synthetic benchmarks and real inference workloads.
  • Real Workload Analysis & Fabric Scalability
  • Run end-to-end inference model workloads on target hardware to capture real-world performance beyond synthetic benchmarks, supporting architecture decisions and customer-facing demonstrations.
  • Evaluate fabric performance as inference cluster size scales from 16 to 32 GPUs and beyond, identifying bottlenecks and building performance scaling models for state-of-the-art AI workloads.
  • Design and execute head-to-head performance comparisons against competing fabric switch solutions to produce data-driven differentiation evidence.
  • Test Infrastructure & Automation
  • Design, build, and maintain automated lab infrastructure including test execution pipelines, traffic generation tooling, and data collection and reporting systems.
  • Enable repeatable, high-quality, and scalable performance measurements across all hardware configurations, reducing manual effort and accelerating the test cycle.
  • Share infrastructure and playbooks with the Product Applications team to accelerate customer application development and issue resolution.
  • Cross-Functional Impact & Innovation
  • Partner closely with ASIC architecture, firmware, software, Product Definition, Product Applications, and Product Marketing teams to communicate findings, influence design decisions, and resolve performance-impacting issues.
  • Serve as a key technical resource in the early evaluation of new fabric architectures, interconnect technologies (UALink, PCIe Gen 6/Gen 7, Ethernet, UEC), and AI/ML communication paradigms.
  • Provide performance data, analysis, and live benchmark support for key customer engagements and industry events; produce clear, audience-appropriate performance reports, technical briefs, and marketing collateral, and maintain living documentation in Confluence.

Basic Qualifications

  • Bachelor's degree in Computer Engineering, Computer Science, Electrical Engineering, or a related technical field.
  • We welcome both recent graduates with strong, directly relevant project, research, or internship experience and candidates with 2–5 years of industry experience in performance or systems engineering.
  • Hands-on experience running AI/ML workloads on GPU clusters — including benchmarking, performance analysis, and fine-tuning of workloads across clusters of GPUs or accelerators.
  • This can come from industry, research, or substantial academic projects.
  • Demonstrated ability to debug and root-cause system-level performance issues across hardware, firmware, software, and network boundaries.
  • Excellent fundamental knowledge of compute algorithms, parallel algorithms, and AI/ML algorithms and workloads.
  • Strong working knowledge of computer systems, GPU systems, and datacenter networking — including PCIe and Ethernet fundamentals.
  • Working knowledge of GPU and CPU software stacks (e.g., CUDA, MPI, collective communication libraries, drivers, and OS-level performance tooling).
  • Proficiency in scripting and automation (e.g., Python) to build test pipelines and analyze large performance datasets.

Preferred Qualifications

  • MS or PhD in Computer Engineering, Computer Science, Electrical Engineering, or a related field.
  • Experience with scale-up fabrics and next-generation interconnects such as UALink, PCIe Gen 6/Gen 7, Ethernet, or UEC.
  • Deep understanding of modern inference and training workloads (LLMs, MoE, recommender systems) and their communication patterns.
  • Experience developing roofline models and competitive performance analyses for switching, networking, or accelerator silicon.
  • Excellent written and verbal communication skills, with the ability to translate deep technical findings into concise executive summaries and customer-facing narratives.

Pay

Salary range is $135,000 to $170,000 depending on experience, level, and business need. This role may be eligible for discretionary bonus, incentives and benefits.

Similar jobs

Quality Engineer

Corning IncorporatedRochester, NY· 3 days ago
Engineeringapply on bit.ly

Quality Engineer

MANITOU GroupMadison, SD· 3 days ago
$60k–$85k/yrapply on careers.manitou-group.com

Quality Engineer

Bastion Technologies, Inc.New Orleans, LA· 3 wk ago
Engineeringapply on bastiontechnologies.applicantpro.com

Quality Engineer

Amsted AutomotiveTaylor, MI· 1 mo ago
Quality Assuranceapply on ars2.equest.com

Quality Engineer

Orion IndustriesAuburn, WA· 1 mo ago
Quality Assuranceapply on recruiting.paylocity.com

Quality Engineer

ActalentJeffersonville, KY· 1 mo ago
Quality Assurance$43.27–$45.67/hrapply on ars2.equest.com