Software Engineer, Systems ML Engineering
Meta · Bellevue, WA · 2 wk ago
Information Technology$184k–$257k/yrFull-time
Responsibilities
- Design and implement scalable ML systems infrastructure components, including distributed training frameworks, model serving pipelines, and ML platform tooling used across Meta's production AI workloads
- Lead technical design and architecture for major initiatives in the ML systems stack, evaluating trade-offs across performance, reliability, and engineering complexity
- Identify and resolve performance bottlenecks in distributed ML training and inference systems through instrumentation, profiling, and targeted optimization
- Define and drive service level objectives for ML infrastructure services, building dashboards, alerting, and runbooks to reduce mean time to mitigation during incidents
- Collaborate with machine learning researchers, product engineers, and infrastructure teams to translate model development requirements into robust, production-grade systems
- Leverage AI-assisted development workflows to accelerate implementation, code review, and system analysis, applying sound judgment on when to rely on AI tooling versus deep domain expertise
- Mentor other engineers on ML systems best practices, distributed computing patterns, and engineering craft, including AI-native development workflows
- Drive adoption of engineering standards across the team, including testing strategies, staged rollout practices using feature flagging and experimentation frameworks, and proactive monitoring
- Contribute to roadmap definition and stakeholder alignment for multi-quarter ML infrastructure investments, communicating technical options and trade-offs to both engineering and cross-functional audiences
Requirements
- Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
- 8+ years of experience in software engineering with a focus on systems software, distributed computing, or ML infrastructure
- Experience designing and implementing large-scale distributed systems, including components such as training orchestration, model serving, or data pipeline infrastructure
- Experience with performance analysis and optimization of compute-intensive or distributed workloads, including profiling, benchmarking, and bottleneck identification
- Experience leading end-to-end delivery of complex technical projects, including cross-team coordination, milestone planning, and risk mitigation
- Experience with C++, Python, or equivalent systems programming languages applied to production ML or infrastructure systems
Qualifications
- Experience contributing to or maintaining open-source ML systems or distributed computing projects
- Experience building or operating ML platform services including experiment tracking, model registries, feature stores, or inference serving infrastructure
- Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
- Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
- Experience with ML frameworks such as PyTorch, including distributed training paradigms such as data parallelism, model parallelism, or pipeline parallelism
- Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
- Experience with GPU computing, CUDA programming, or accelerator-aware systems optimization for large-scale AI workloads