Principal Machine Learning Engineer
Haystack · United States · 1 mo ago
RemoteRemoteEngineeringFull-time
About the role
Design and evolve critical large-scale ML systems across training, inference, evaluation, and infrastructure. Set technical standards and shape ML system development across the organization.
Responsibilities
- Architect and build large-scale ML systems spanning data, training, evaluation, inference, and deployment
- Design reproducible, high-performance training pipelines on GPU infrastructure
- Implement evaluation pipelines for performance, robustness, safety, and bias
- Own production deployment and collaborate with application engineering for ML system integration
Requirements
- Strong background in deep learning and transformer-based architectures
- Artificial Intelligence (AI) experience is required
- Hands-on experience training, fine-tuning, or deploying large-scale ML models in production
- Proficiency with at least one modern ML framework (e.g., PyTorch, JAX)
- Experience with distributed training and inference frameworks
- Strong software engineering fundamentals and experience with GPU optimization
Benefits
- Comprehensive medical, dental, and vision insurance
- Savings plan options
- Generous paid time off (PTO)
- Opportunity to work 100% remotely