Generative AI - ML System Engineering
About the role
We are looking for Machine Learning Systems Engineers to help build the world's largest end-to-end 3D native machine learning systems. You will work on building our end to end ML framework dedicated for 3D, from pretraining, to finetuning, inferencing, etc.
Responsibilities
- Work closely with researchers to co-design the next frontier of 3D & Spatial AI.
- Build and debug on top of modern PyTorch, for maximum parallelism and efficiency.
- Build clean and intuitive training infrastructure for our in-house foundational models.
- Identify bottlenecks and optimize for high throughput & efficient distributed model training across hundreds to thousands of GPUs.
- Implement and maintain 3D specific custom operators in Triton or CUDA.
- Implement and maintain novel data-loading framework and libraries.
- Build efficient inference endpoints with complex multi-stage model pipelines.
- Optimize models through compilation, fusion, quantization, etc.
Requirements
Experience in machine learning or high performance graphics.
Strong practical understanding of at least one machine learning framework (e.g. PyTorch, JAX).
Strong ability to write beautiful and maintainable code in Python and/or C++.
Ability to learn fast and dive into new concepts or complex codebases.
Performance and efficiency oriented mindset, with a strong interest in the tiniest detail.
Strong communication skills for working in a globally distributed team.
Qualifications
A strong passion to navigate through the PyTorch internals, with hands-on experience in areas like torch.compile, fully_shard (FSDP2) APIs.
Experience with building Triton kernels.
Familiarity with modern parallelization techniques: DP, TP, CP, PP, zero redundancy optimizers, etc.
Experience with diffusion models in 3D or video.
Experience with low precision bf16 or fp8 training.
Skills
A strong passion to navigate through the PyTorch internals, with hands-on experience in areas like torch.compile, fully_shard (FSDP2) APIs.
Experience with building Triton kernels.
Familiarity with modern parallelization techniques: DP, TP, CP, PP, zero redundancy optimizers, etc.
Experience with diffusion models in 3D or video.
Experience with low precision bf16 or fp8 training.
Benefits
Trusted by Meta, Square Enix, Deepmind and more, Meshy is redefining 3D creation with generative AI. We empower artists, designers, engineers, hobbyists, and makers to bring immersive worlds, characters, and experiences to reality in minutes instead of months.
We have a culture of directness and truthfulness, therefore we value constructive criticism. Being direct and truthful is the most sincere form of trust and care.
We have a keen eye for quality and aesthetics. Our products are not just functional but also beautiful. The same aesthetics permeate through our culture, our code and are the same: functional and beautiful.
Pay
The base salary range for this position is $175,000 – $300,000 per year.
Schedule
This is a fully in-office position where you will be an integral part of our daily operations.