Machine Learning Engineer — Inference Optimization
Jobgether · United States · 1 wk ago
RemoteRemoteEngineeringFull-time
Accountabilities
- Owes the performance and scalability of machine learning models in production.
- Combines deep ML expertise, systems engineering, and performance analysis to deliver faster, more efficient AI experiences.
- Optimizes machine learning inference systems to improve latency, throughput, scalability, and operational cost.
- Profile and identify bottlenecks across GPU and CPU inference pipelines, including memory usage, kernels, batching strategies, and data flow.
- Implement advanced optimization techniques such as quantization, KV-cache optimization, speculative decoding, batching, streaming, and model simplification.
- Collaborate with research engineers to productionize new model architectures and translate experimental results into reliable systems.
- Build, improve, and maintain inference-serving infrastructure using modern frameworks, custom runtimes, or specialized serving solutions.
- Benchmark model performance across different hardware environments, including GPUs, CPUs, and cloud-based systems.
- Improve system reliability, monitoring, observability, and cost efficiency under real production workloads.
- Contribute to engineering practices that improve the quality, scalability, and maintainability of ML infrastructure.
Requirements
- Tech-savvy machine learning engineer with experience optimizing production inference systems.
- Passion for high-performance AI engineering and complex technical problem-solving.
- Strong professional experience in ML inference optimization, high-performance machine learning systems, or related areas.
- Deep understanding of machine learning fundamentals, including neural network architectures, attention mechanisms, memory optimization, and compute graphs.
- Hands-on experience with PyTorch or similar deep learning frameworks and deploying models into production environments.
- Experience with GPU performance optimization, including technologies such as CUDA, ROCm, Triton, or kernel-level tuning.
- Proven experience scaling inference systems for real users beyond research prototypes or benchmarks.
- Strong programming skills and ability to work across machine learning and systems engineering domains.
- Able to operate effectively in fast-paced environments with ownership, autonomy, and evolving priorities.
- Experience with inference frameworks such as TensorRT, ONNX Runtime, vLLM, or Triton is a plus.
- Familiarity with large language models, long-context inference, distributed systems, low-latency services, or hardware optimization is considered an advantage.
- Contributions to open-source ML systems or inference tooling are a plus.
Benefits
- Competitive compensation package with meaningful equity participation.
- Opportunity to work on performance-critical AI systems with direct product impact.
- Closer collaboration with research, infrastructure, and product teams.
- Opportunity to work on advanced machine learning technologies and real-world AI applications.
- Engineering-focused culture that values technical excellence, experimentation, and quality.
- Flexible remote work environment.
- Opportunity to contribute to the growth of an innovative AI-focused organization.