Senior Machine Learning Engineer, Jockey Core
Seoulstart · Range Line, IN · 4 days ago
Information TechnologyFull-time
Seoul, South Korea
About the role
Build benchmarks and load tests replaying agent real traffic, measuring TTFT and inter-token latency separately. Apply inference optimization techniques (quantization, batching, disaggregated prefill/decode, speculative decoding) to hit cost and latency targets. Build cost models and drive production hardening through autoscaling, capacity planning, observability, failover, and rollback-safe rollouts.
Requirements
- Significant experience serving and optimizing large-scale LLM inference in production (vLLM, TensorRT-LLM, SGLang, or similar).
- Experience with batching, scheduling, quantization, disaggregated prefill/decode, and speculative decoding techniques.
- Experience designing and operating large-scale distributed systems in high-performance GPU environments.
Nice to have
- Experience contributing to or customizing internals of LLM inference servers (vLLM, TensorRT-LLM, SGLang, or similar).
- Understanding of serving characteristics of compressed (pruned or quantized) models or reasoning or agentic LLMs.
- Experience designing multi-region or multi-cluster serving infrastructure or large-scale GPU capacity planning.
Context
Foreign-invested AI company (San Francisco HQ, Seoul R&D office) with strong visa sponsorship track record; role is fully English-language, part of global Cognition Models team, likely F-series visa pathway (F-2 / F-4 / F-5 / F-6 holders welcome). Korean language not required.