AI Infrastructure Engineer
About the role
We are seeking an experienced AI Infrastructure Engineer to join our AI Incubation team. You will be focused on building and optimizing large-scale training infrastructure for Large Language Models (LLMs).
Responsibilities
Designing and developing scalable AI infrastructure solutions for training and deploying large language models
Building and optimizing distributed training platforms using cutting-edge technologies
Implementing and maintaining containerized AI environments using Docker and Kubernetes
Optimizing CUDA kernels for maximum GPU utilization and performance
Developing platform software to support AI/ML workflows
Collaborating with AI researchers to implement efficient training and inference pipelines
What We’re Looking For
Have a bachelor's degree in Computer Science, Engineering, AI, Machine Learning, Distributed System or related field
5+ years of software engineering experience with focus on infrastructure and systems
Expertise in GPU programming and CUDA optimization
Experience with container technologies (Docker, Kubernetes), distributed systems and cloud computing
Experience building large-scale distributed systems and optimizing neural network performance
Programming skills in Python, C++, and CUDA, with deep learning frameworks (PyTorch, Transformers)
Pay
Minimum Salary Range or On Target Earnings: $151,800.00 - $332,200.00
Ways of Working
Our structured hybrid approach is centered around our offices and remote work environments. The work style of each role, Hybrid, Remote, or In-Person is indicated in the job description/posting.
Benefits
As part of our award-winning workplace culture and commitment to delivering happiness, our benefits program offers a variety of perks, benefits, and options to help employees maintain their physical, mental, emotional, and financial health; support work-life balance; and contribute to their community in meaningful ways.