Member of Technical Staff (GPU Performance Engineer)
About the role
You will design and implement improvements to our training infrastructure and directly contribute to technical decisions that optimize performance of our models. You will also work on post-training processes, including reinforcement learning and fine-tuning. Furthermore, you will contribute to improving the efficiency and scalability of our model serving infrastructure.
Responsibilities
- Design and implement improvements to training infrastructure
- Contribute to technical decisions that optimize model performance
- Work on post-training processes such as reinforcement learning and fine-tuning
- Improve efficiency and scalability of model serving infrastructure
Requirements
- Strong engineering skills with fluency in Python and PyTorch (or other frameworks)
- Proven experience implementing and training large deep learning models
- Experience writing and debugging low-level GPU code (CUDA, C++)
- Experience scaling up GPU jobs using large-scale compute clusters (e.g., Slurm or Kubernetes)
- Demonstrated ability to analyze and optimize the performance of GPU-accelerated workloads, including profiling, identifying bottlenecks, and implementing performance tuning techniques
Qualifications
Reka's Mission: Reka's mission is to build useful multimodal artificial intelligence and use it to empower organizations and businesses. We are a globally distributed foundation model startup, headquartered in the San Francisco Bay Area, California. Embracing a remote-first approach, our team brings together top talent from around the world.
Skills
- Python
- PyTorch (or other frameworks)
- CUDA, C++
- Large-scale compute clusters (Slurm, Kubernetes)
- Performance analysis and optimization
Benefits
- Five weeks of paid leave
- Comprehensive healthcare benefits (vision and dental)
- Additional perks supporting well-being
Pay
TBD
Schedule
TBD
Benefits
- Visa support (H1B and OPT transfers)
Reka
Reka is an elite team of top-tier engineers, researchers, and operators from renowned organizations like Google DeepMind, Facebook AI Research (FAIR), and successful startups. We drive innovation in AI technology and leverage cutting-edge infrastructure to train state-of-the-art models. Our inclusive and open culture values diverse perspectives and fosters creativity. We offer generous benefits, including five weeks of paid leave, comprehensive healthcare benefits, and additional perks to support your well-being.