Senior Staff Engineer, AI Software
Samsung Semiconductor · San Jose, CA · 1 mo ago
Engineering$189k–$301k/yrFull-time
About the role
The AGI (Artificial General Intelligence) Computing Lab is dedicated to solving the complex system-level challenges posed by the growing demands of future AI/ML workloads. Our team is committed to designing and developing scalable platforms that can effectively handle the computational and memory requirements of these workloads while minimizing energy consumption and maximizing performance.
Responsibilities
- Lead the co-design of software and hardware solutions that optimize AI model inference performance, with a focus on overcoming memory bottlenecks.
- Analyze and optimize LLM and agentic AI workloads across the full software stack, identifying opportunities for hardware-aware acceleration.
- Profile and characterize model execution to expose memory wall limitations and guide architectural decisions for HBM and memory-centric compute.
- Collaborate with hardware teams to influence memory architecture, acceleration strategies, and compute placement based on real workload behavior.
- Develop, optimize, and benchmark inference and serving solutions using frameworks such as PyTorch and vLLM.
- Define best practices and provide technical mentorship across software–hardware co-design efforts.
Requirements
- Bachelor’s with 15+ years, or Master’s with 13+ years, or PhD's with 10+ years of industry experience.
- Strong experience writing high-performance AI framework software development for GPUs or other accelerators.
- Strong, end-to-end understanding of the AI infrastructure, AI software stack, from model definition through deployment and serving.
- Solid understanding of LLM model architectures and workflows, including modern transformer-based designs.
- Solid understanding of agentic AI architecture and workflows.
- Hands-on expertise with the PyTorch framework.
- PRACTICAL experience with the vLLM for high-throughput model inference and serving.
- Strong knowledge of the memory wall problem and its impact on AI system performance.
- Strong knowledge of memory architecture, including High Bandwidth Memory (HBM), and familiarity with memory-centric acceleration and compute approaches.
- Proficiency working in a Linux development environment.
- Solid command of development tooling, including agentic coding, GitHub and Jira.
Qualifications
- Bachelor’s degree in Computer Science, Electrical Engineering, or related field.
- Master’s degree in Computer Science, Electrical Engineering, or related field.
- PhD in Computer Science, Electrical Engineering, or related field.
Skills
- Experience with high-performance AI framework software development for GPUs or other accelerators.
- Understanding of LLM model architectures and workflows, including modern transformer-based designs.
- Understanding of agentic AI architecture and workflows.
- Hands-on expertise with the PyTorch framework.
- Practical experience with the vLLM for high-throughput model inference and serving.
- Knowledge of the memory wall problem and its impact on AI system performance.
- Knowledge of memory architecture, including High Bandwidth Memory (HBM), and familiarity with memory-centric acceleration and compute approaches.
- Proficiency working in a Linux development environment.
- Command of development tooling, including agentic coding, GitHub and Jira.
Benefits
We offer a comprehensive benefits package that includes:
- Medical/Dental/Vision coverage.
- 4+ weeks of paid time off a year, plus holidays and sick leave.
- Support for fertility care or adoption, medical travel, and virtual vet care for your fur babies.
- On-demand mental health resources and confidential therapy sessions.
- EatWell and MoveWell programs with onsite Café and gym, plus virtual classes.
- A flexible work environment to help you find the right balance for you.
Pay
$189,000—$301,000 USD
Schedule
Daily onsite presence at our San Jose, CA office / U.S. headquarters in alignment with our Flexible Work policy.