GPU Systems Engineer – HPC / Parallel Computing
Vast.ai’s cloud powers AI projects and businesses all over the world. We are democratizing and decentralizing AI computing—reshaping our future for the benefit of humanity. We are a growing and highly motivated team dedicated to an ambitious technical plan. Our structure is flat, our ambitions are out-sized, and leadership is earned by shipping excellence. We seek engineers with strong intrinsic drive, a true passion for advancing the state of the art, and a mix of architecture, coding, and communication skills.
Location: On-site at our office in San Francisco or Westwood, Los Angeles.
About the Role
We’re looking for a systems engineer with HPC or parallel programming experience to help scale AI inference. You’ll leverage your knowledge of high-performance systems to optimize GPU performance at the bleeding edge of AI.
Tech Stack: CUDA/C++, GPGPU, Python, Linux
Responsibilities
- Design and optimize GPU kernels and tensor libraries
- Translate HPC techniques into scalable AI inference solutions
- Evaluate emerging architectures and resource management approaches
- Collaborate with technical leadership to improve GPU infrastructure efficiency
Requirements
- Advanced C++ (C++17/20 preferred)
- Expertise with at least one parallel framework (CUDA, HIP, SYCL, OpenCL, OpenACC, or similar)
- Strong background in systems optimization and HPC performance tooling
- Familiarity with distributed training/inference frameworks (bonus)
Benefits
- Comprehensive health, dental, vision, and life insurance
- 401(k) with company match
- Meaningful early-stage equity
- Onsite meals, snacks, and close collaboration with founders/tech leaders
- Ambitious, fast-paced startup culture where initiative is rewarded
Pay
Compensation Range: $160K - $320K