Jobs · Engineering · California

AI Inference Engineer

Quadric · Burlingame, CA · 3 wk ago
HybridEngineering$110k–$270k/yrFull-time

Quadric has created an innovative general purpose neural processing unit (GPNPU) architecture. Quadric's co-optimized software and hardware is targeted to run neural network (NN) inference workloads in a wide variety of edge and endpoint devices, ranging from battery-operated smart-sensor systems to high-performance automotive or autonomous vehicle systems. Unlike other NPUs or neural network accelerators in the industry today that can only accelerate a portion of a machine learning graph, the Quadric GPNPU executes both NN graph code and conventional C++ DSP and control code.

About the role

The AI Inference Engineer at Quadric is the key bridge between the world of AI/LLM models and Quadric’s unique platforms. This role involves porting AI models to the Quadric platform, optimizing model deployment for efficient inference, and profiling and benchmarking model performance. This senior technical role demands deep knowledge of AI model algorithms, system architecture, and AI toolchains/frameworks. This California Bay Area-based role follows a hybrid schedule, with at least two in-office days per week at our Burlingame office, the ability to commute regularly, and occasional additional onsite days as needed based on team and business priorities.

Responsibilities

  • Quantize, prune, and convert models for deployment
  • Port models to Quadric platform using Quadric toolchain
  • Optimize inference deployment for latency and speed
  • Benchmark and profile model performance and accuracy
  • Collaborate across related areas of the AI inference stack to support team and business priorities
  • Develop tools to scale and speed up deployment
  • Make improvements to SDK and runtime
  • Provide technical support and documentation to customers and the developer community

Requirements

  • Bachelor's or Master's in Computer Science and/or Electrical Engineering
  • 5+ years of experience in AI/LLM model inference and deployment frameworks/tools
  • Experience with model quantization (PTQ, QAT) and tools
  • Experience with model accuracy measures
  • Experience with model inference performance profiling
  • Experience with at least one of the following frameworks: ONNX Runtime, PyTorch, vLLM, Hugging Face Transformers, Neural Compressor, llama.cpp
  • Proficiency in C/C++ and Python
  • Demonstrated capability in problem-solving, debugging, and communication

Benefits

  • Competitive salary and meaningful equity
  • Medical, dental, and vision plans starting on day one
  • 401(k) retirement plan
  • Flexible paid time off (unlimited, non-accrual) to support work-life balance
  • Company-provided lunches and a stocked kitchen when working in-office
  • Convenient office location within walking distance of the Caltrain station
  • Support for commuting, including monthly parking or Caltrain passes
  • Downtown Burlingame office location, close to shops, cafes, and local amenities
  • A politics-free, highly collaborative environment where talented people can do their best work
  • Opportunity to build long-term career relationships in a company that values strong personal connections alongside professional excellence

Pay

The base salary range for this position is $110,000 to $270,000. The actual base salary offered will depend on factors such as the specific level of the role, years and depth of relevant experience, technical skills and competencies, the criticality of the role to the business, internal equity, and work location. In addition to base salary, this role is eligible for equity and a discretionary annual performance bonus.

Similar jobs

AI Engineer

AI FundMountain View, CA· 1 mo ago
Engineering$130k–$175k/yrapply on jobs.lever.co

AI Engineer

MillenniumNew York, NY· 1 mo ago
Information Technology$175k–$250k/yrapply on mlp.eightfold.ai

AI Engineer

PepsiCoPurchase, NY· 1 mo ago
Engineeringapply on pepsicojobs.com

AI Engineer

Stanford UniversityRedwood City, CA· 1 mo ago
Engineering$170k–$195k/yrapply on careersearch.stanford.edu

AI Engineer

P-1 AISan Francisco Bay Area· 1 mo ago
Engineeringapply on jobs.ashbyhq.com

AI Engineer

Rockland TrustMassachusetts, United States· 3 mo ago
Engineeringapply on eehb.fa.us2.oraclecloud.com