Jobs · Engineering

Machine Learning Engineer — Inference Optimization

Jobgether · United States · 1 wk ago
RemoteRemoteEngineeringFull-time

Accountabilities

  • Owes the performance and scalability of machine learning models in production.
  • Combines deep ML expertise, systems engineering, and performance analysis to deliver faster, more efficient AI experiences.
  • Optimizes machine learning inference systems to improve latency, throughput, scalability, and operational cost.
  • Profile and identify bottlenecks across GPU and CPU inference pipelines, including memory usage, kernels, batching strategies, and data flow.
  • Implement advanced optimization techniques such as quantization, KV-cache optimization, speculative decoding, batching, streaming, and model simplification.
  • Collaborate with research engineers to productionize new model architectures and translate experimental results into reliable systems.
  • Build, improve, and maintain inference-serving infrastructure using modern frameworks, custom runtimes, or specialized serving solutions.
  • Benchmark model performance across different hardware environments, including GPUs, CPUs, and cloud-based systems.
  • Improve system reliability, monitoring, observability, and cost efficiency under real production workloads.
  • Contribute to engineering practices that improve the quality, scalability, and maintainability of ML infrastructure.

Requirements

  • Tech-savvy machine learning engineer with experience optimizing production inference systems.
  • Passion for high-performance AI engineering and complex technical problem-solving.
  • Strong professional experience in ML inference optimization, high-performance machine learning systems, or related areas.
  • Deep understanding of machine learning fundamentals, including neural network architectures, attention mechanisms, memory optimization, and compute graphs.
  • Hands-on experience with PyTorch or similar deep learning frameworks and deploying models into production environments.
  • Experience with GPU performance optimization, including technologies such as CUDA, ROCm, Triton, or kernel-level tuning.
  • Proven experience scaling inference systems for real users beyond research prototypes or benchmarks.
  • Strong programming skills and ability to work across machine learning and systems engineering domains.
  • Able to operate effectively in fast-paced environments with ownership, autonomy, and evolving priorities.
  • Experience with inference frameworks such as TensorRT, ONNX Runtime, vLLM, or Triton is a plus.
  • Familiarity with large language models, long-context inference, distributed systems, low-latency services, or hardware optimization is considered an advantage.
  • Contributions to open-source ML systems or inference tooling are a plus.

Benefits

  • Competitive compensation package with meaningful equity participation.
  • Opportunity to work on performance-critical AI systems with direct product impact.
  • Closer collaboration with research, infrastructure, and product teams.
  • Opportunity to work on advanced machine learning technologies and real-world AI applications.
  • Engineering-focused culture that values technical excellence, experimentation, and quality.
  • Flexible remote work environment.
  • Opportunity to contribute to the growth of an innovative AI-focused organization.

Similar jobs