Lead AI Infrastructure Engineer
About the Team
Our team is at the core of Avride's self-driving stack. We build the base infrastructure layer that powers all autopilot code. It includes a C++ framework for implementing autonomy components, execution graph building and optimization systems, as well as runtimes that execute those graphs, both onboard and in simulation. The vast part of the execution graph is implemented as a chain of neural network operations. The onboard mode relies on stable latencies of the inference of those networks, while in simulation we also optimize throughput at scale.
About the Role
We’re looking for a software engineer with a leadership mindset and deep ML infrastructure experience. You will decide and influence the ML infrastructure layer across the company. The biggest challenge we’re facing at the moment is the effectiveness of GPU inference - both for onboard applications with near real-time guarantees and for offboard cases that target high throughput and deterministic execution. It is the first priority within this role.
Responsibilities
- Take on the GPU inference framework, focusing on performance
- Assume responsibility and ownership for broader ML infrastructure scattered across ML pipelines
- Close collaboration with the applied ML team responsible for defining the neural model's architecture
Requirements
- Experience with PyTorch
- Understanding of how GPUs work
- Experience in diagnosing and resolving performance issues
- Strong record of building infrastructure including distributed systems
- 5+ years of experience with C++
- Programming experience in multi-threaded environments - multiple processes, threads, timers, and interrupts
Additional Information
Candidates are required to be authorized to work in the U.S. The employer is not offering relocation sponsorship, and remote work options are not available.