Applied AI Engineer, Kernel Performance
About the role
Own the system that turns new model architectures into verified, production-ready kernels and model mappings.
Build agents that understand Etched hardware, design experiments, generate implementations, compile and profile them, diagnose bottlenecks, and iterate with our teams, to the limits of model autonomy.
Design evals covering correctness, numerical stability, latency and efficiency.
Turn profiler traces, simulation, hardware counters, and expert judgment into structured signals models can learn from.
Curate proprietary datasets and optimization memory from complete trajectories, expert demonstrations, counterexamples, and production outcomes.
Build fast, reproducible experiment infrastructure and observability so experiments remain interpretable, trustworthy, and high-throughput.
Ship model-generated improvements to production and quantify their impact on end-to-end system performance.
Partner deeply with other architecture teams to shape new abstractions and Etched’s hardware-software roadmap.
Continuously evaluate new model releases and deploy the best for each stage of the optimization loop.
Responsibilities
- Own the system that turns new model architectures into verified, production-ready kernels and model mappings.
- Build agents that understand Etched hardware, design experiments, generate implementations, compile and profile them, diagnose bottlenecks, and iterate with our teams, to the limits of model autonomy.
- Design evals covering correctness, numerical stability, latency and efficiency.
- Turn profiler traces, simulation, hardware counters, and expert judgment into structured signals models can learn from.
- Curate proprietary datasets and optimization memory from complete trajectories, expert demonstrations, counterexamples, and production outcomes.
- Build fast, reproducible experiment infrastructure and observability so experiments remain interpretable, trustworthy, and high-throughput.
- Ship model-generated improvements to production and quantify their impact on end-to-end system performance.
- Partner deeply with other architecture teams to shape new abstractions and Etched’s hardware-software roadmap.
- Continuously evaluate new model releases and deploy the best for each stage of the optimization loop.
Requirements
A track record of solving hard problems across stacks and domains — you enjoy being dropped into unfamiliar territory and figuring it out
Comfort with both Python and low-level code: you can read it, modify it, debug it, and direct AI to write it well. We do not care whether you write code from scratch — we care whether you ship things that work
Kernel experience: you've written or tuned kernels and can explain the mechanisms and performance impact of optimizations you’ve shipped
Fluency using AI to learn and ramp on new problems — agentic coding tools, deep research, and frontier models are how you work, not an add-on
Moving fluidly between research exploration, agentic experimentation, low-level debugging, and production execution
Strong Candidates May Also Have Experience With
First principles thinking on accelerator performance: memory hierarchy, data movement, parallelism, synchronization, and low-precision computation
Hands-on experience building and shipping LLM-based agents or AI tooling that real users depend on in production environments (beyond calling an API — context engineering, tool integration, orchestration, failure analysis)
An eval-driven mindset: you measure whether AI systems work before scaling them
Fine-tuning or post-training, RAG over proprietary data, and/or multi-agent orchestration
High agency and comfort with ambiguity — you find the real problem to solve
Qualifications
Experience with AI and machine learning, particularly in the context of large language models and transformer architectures.
Experience with compiler optimization and low-level programming.
Experience with hardware design and performance tuning.
Experience with experimental design and evaluation methods.
Skills
Experience with AI and machine learning, particularly in the context of large language models and transformer architectures.
Experience with compiler optimization and low-level programming.
Experience with hardware design and performance tuning.
Experience with experimental design and evaluation methods.
Benefits
- Medical, dental, and vision packages with generous premium coverage
- $500 per month credit for waiving medical benefits
- Housing subsidy of $2k per month for those living within walking distance of the office
- Relocation support for those moving to San Jose (Santana Row)
- Variety of wellness benefits covering fitness, mental health, and more
- Daily lunch and dinner in our office
- Unlimited compute budget subject to ROI justification
Pay
Base Compensation Range $150,000 – $225,000
Schedule
Full-time, in-person position in San Jose, CA