Accelerator Compiler and Tool Chain Lead
Santa Clara, CA • Onsite • Full Time
About Velaura AI
Velaura AI is a semiconductor and technology company that provides patented ultra-low-power silicon design technology, IP, toolflows, and custom chiplet solutions for AI compute platforms. Its customers include hyperscaler and XPU companies seeking reduced power consumption and higher compute efficiency. The company also builds Teraflux Bitcoin mining products, including air-, hydro-, and immersion-cooled miners, ASICs, modular containers, miner firmware, fleet-management software, and enterprise customer support.
About The Role
You will lead the compiler and model-lowering stack for an AI accelerator. You will own model ingestion, graph lowering, compiler IR, optimization passes, quantization integration, code generation, diagnostics, verification, and regression strategy while leading a compiler and ML systems engineering team.
Responsibilities
- Lead architecture and development of the AI accelerator compiler stack
- Own model ingestion and graph lowering from ML frameworks and exchange formats
- Define operator coverage, lowering rules, graph transformations, fusion, partitioning, and fallback behavior
- Develop compiler optimization passes for tensor layout, tiling, memory movement, mixed precision, operator fusion, and hardware scheduling
- Define executable artifact formats, metadata, memory planning requirements, profiling hooks, and runtime constraints
- Integrate quantization compilation
- Build compiler diagnostics for unsupported operators, shape constraints, graph rewrites, quantization issues, and performance bottlenecks
- Establish compiler verification and regression strategies
- Hire, mentor, and lead compiler and ML systems engineers
Requirements
- Deep experience with compiler development, ML graph compilers, or accelerator code generation
- Strong understanding of ML model formats, graph IRs, operator lowering, tensor layouts, quantization, and runtime/compiler interfaces
- Strong C++ and Python programming skills
- Experience with MLIR, LLVM, TVM, XLA, IREE, Glow, TensorRT-like systems, OpenVINO-like systems, or equivalent
- Understanding of correctness risks in compiler optimizations, graph rewrites, mixed precision, operator fusion, and hardware-specific lowering
- Ability to collaborate with hardware, firmware, runtime, model-integration, and SQA engineers
- Experience leading technical teams or major architecture areas
Skills
- AI Accelerator
- C#
- Code Generation
- Compiler
- Compiler Diagnostics
- Graph IR
- Graph Partitioning
- IREE
- LLVM
- Mixed Precision
- ML Compiler
- MLIR
- NPU
- ONNX
- Operator Lowering
- Python
- PyTorch
- Quantization
- Runtime
- TensorFlow Lite
- Tensor Layout
- TVM
- XLA
Benefits
- Performance-based incentives
- Equity participation
- Medical coverage
- Dental coverage
- Vision coverage
- Paid time off
- Flexible work arrangements
Pay
$200,000 – $500,000