Jobs · Engineering · California

AI Inference Systems Architect

Ludwig Computing · San Francisco Bay Area · 4 days ago
HybridEngineeringFull-time

About Us

At Ludwig Computing, we are solving the energy efficiency problem of intelligent compute. Our novel co-designed approach is optimized to deliver radical improvements in energy efficiency and performance across a wide range of AI workloads. We are building a future where high-performance computing is powered by leaner, smarter, and extremely efficient hardware and software platforms. Join us at the ground floor as we build the future of intelligent compute.

Role Description

We are looking for a highly experienced, hands-on ML Systems Architect to define the next generation of large language model (LLM) inference architecture. This role sits at the intersection of hardware and software: you will reason about the entire pipeline — from a prompt (or batch of prompts) arriving at the system to the generation of each output token — and use that understanding to guide architecture decisions across compute, memory, and interconnect. You will be a key voice in shaping what the next macro-architecture for inference should look like, anticipating how emerging model trends (e.g., mixture-of-experts, longer context windows, new attention mechanisms) will reshape system requirements.

What You'll Do

  • Hardware–Software Co-Design & Macro-Architecture
    • Analyze the full LLM inference pipeline end to end, from prompt ingestion to token generation, to identify where architectural investment matters most.
    • Define next-generation macro-architecture directions by anticipating how emerging model families and workload trends will change system requirements.
    • Partner with hardware and software teams to co-design solutions, ensuring architectural decisions reflect real workload and model behavior rather than either discipline in isolation.
  • Performance Modeling & Metric Trade-offs
    • Model inference performance as a function of the key system metrics: on-chip SRAM capacity, off-chip memory bandwidth, compute throughput, interconnect bandwidth/latency, and other relevant parameters.
    • Understand how adjusting each metric shifts overall system performance, and use that understanding to guide design trade-offs and optimization priorities.
  • Bottleneck Diagnosis
    • Identify the dominant bottleneck for a given workload and prescribe the right fix, for example:
    • Decode bottlenecked by bandwidth → improve memory bandwidth, compression, caching, batching, or weight reuse.
    • Long context bottlenecked by KV cache → change KV layout, memory capacity, interconnect, paging, or attention implementation.
    • Multi-card inference bottlenecked by communication → change partitioning, topology, or parallelism strategy.
    • Low utilization → fix runtime scheduling, batching, operator fusion, or DMA overlap.
  • Multi-XPU Inference Strategy
    • Design and evaluate multi-XPU inference strategies spanning model partitioning, tensor parallelism, and expert parallelism.
    • Determine optimal KV-cache placement and estimate communication volume per token across devices.
    • Define bandwidth and latency requirements for a given topology, and assess when a ring topology is sufficient versus when a switch fabric becomes unavoidable.
  • System-Level Benchmarking
    • Build performance models and benchmarks at the full system level, not just at the individual GPU/accelerator level, to enable apples-to-apples evaluation of SoC architecture alternatives.
    • Use these models to inform build-vs-buy and architecture trade-off decisions early in the design cycle.

What We're Looking For

  • 5–10+ years of experience in ML systems architecture, AI accelerator/XPU design, or inference performance engineering.
  • Hands-on experience with large language models and a strong grasp of transformer-based inference workloads (prefill/decode, attention, KV cache).
  • Demonstrated ability to reason jointly about hardware and software — comfortable moving between silicon-level constraints and runtime/software behavior.
  • Working knowledge of the key system metrics that determine inference performance (SRAM capacity, off-chip bandwidth, compute throughput, interconnect bandwidth/latency, and related parameters) and how to tune them.
  • Practical experience diagnosing system bottlenecks (bandwidth, KV cache/memory capacity, inter-device communication, utilization) and translating diagnoses into concrete architectural fixes.
  • Deep familiarity with multi-XPU/multi-accelerator inference: partitioning strategies, tensor and expert parallelism, KV-cache placement, per-token communication costs, and interconnect topology trade-offs (ring vs. switch fabric).
  • Experience building or using system-level performance models and benchmarks to evaluate SoC architecture alternatives, beyond single-GPU-level analysis.
  • Prior experience at a company developing frontier LLMs and/or a company developing XPUs/AI accelerators as products is highly valued.
  • Strong communication skills and the ability to translate system-level trade-offs into clear architectural recommendations for cross-functional stakeholders.

Nice To Haves

  • Familiarity with emerging model architecture trends (mixture-of-experts, long-context/retrieval-augmented approaches, novel attention mechanisms) and their systems implications.
  • Experience contributing to or influencing SoC/accelerator roadmap decisions.
  • Publications, patents, or public talks related to ML systems, inference performance, or accelerator architecture.

Please direct any questions to apply@ludwigcomputing.com.

Similar jobs