Jobs · Art & Creative · California

Principal Engineer, AI Serving Framework Architect (Software)

Samsung Semiconductor · San Jose, CA · 1 mo ago
Art & Creative$219k–$351k/yrFull-time

About the role

The Architecture Research Lab (ARL) focuses on addressing fundamental system-level bottlenecks in modern AI, particularly in memory capacity/bandwidth and system-scale communication. By leveraging Samsung's world-class memory technologies, ARL explores and defines next-generation AI system architectures that deliver step-function improvements in performance, efficiency, and scalability.

Responsibilities

  • Leading research teams in Korea and proposing technical direction
  • Researching dynamic scheduling methodologies for maximizing AI inference performance in multi-rack scale memory-centric systems, comprised of heterogeneous compute-capable memory and hierarchical memory
  • Investigating methods to accelerate search operations in RAG’s vector DB and AI Agent’s knowledge-graph by leveraging compute-capable memory
  • Studying strategies for optimally placing KVCache and a vector DB in hierarchical memory to minimize frequent SSD accesses and reduce IO stalls
  • Proposing SW design for implementing the derived optimization algorithms on open-source platforms such as vLLM

Requirements

  • PhD in Computer Science or a related field with 10+ years of experience in AI Serving Framework for large-scale computing, with focusing on the AI workloads
  • Led a project to build and optimize a Large Language Model (LLM) Inference Software Stack on a multi-rack scale system to deliver AI Inference services to over 100,000 users
  • Extensive experience in designing AI Inference Software Stacks for heterogeneous devices
  • In-depth understanding of the internal architecture and operation mechanisms of inference engines such as vLLM
  • Proficiency in AI Inference System Profiling and optimization
  • Knowledge and practical experience with future AI workloads, including reasoning models, multi-modal solutions, AI agents, and world models
  • Strong understanding of compute, memory, and networking bottlenecks in AI systems

Qualifications

  • Native or fluent Korean speakers are preferred
  • You’re inclusive, adapting your style to the situation and diverse global norms of our people
  • You approach challenges with curiosity and resilience, seeking data to help build
  • You’re collaborative, building relationships, humbly offering support and openly welcoming approaches
  • You’re innovative and creative, you proactively explore new ideas and adapt quickly to change

Skills

  • PYTORCH
  • Python
  • C++

Benefits

  • Base Pay Range: $219,000—$351,000 USD
  • Comprehensive accommodations for candidates with disabilities, long-term conditions, neurodivergent individuals, or those requiring pregnancy-related support
  • Charitable giving match and frequent opportunities to get involved in the community
  • Time off, holidays, and sick leave
  • Fertility care or adoption stipend, medical travel support, and virtual vet care for fur babies
  • On-demand apps and free confidential therapy sessions for emotional wellness
  • Eat well and stay fit with onsite Café and gym, plus virtual classes
  • Flexible environment to find the right balance for you

Schedule

Daily onsite presence at our San Jose office in alignment with our Flexible Work policy

Pay

$219,000—$351,000 USD

Equal Opportunity Employer

We are committed to fostering an environment where all individuals feel valued and empowered to excel, regardless of race, religion, color, age, disability, sex, gender identity, sexual orientation, ancestry, genetic information, marital status, national origin, political affiliation, or veteran status.

Similar jobs