Sr GenAI Infra Specialist SA, AWS WWSO Startup
About the role
Work directly with the most important and exciting Startup customers in the GenAI model training and inference space, helping them adopt and scale large-scale workloads (e.g., frontier models, models, multi-modal systems, optimization) on AWS
Advise customers on AI infrastructure requirements and trade-offs including GPU/Trainium selection, cluster topology, storage, networking (EFA), and cost optimization for training and inference
Provide deep technical guidance on inference optimization model serving architectures (self-managed on EKS, SageMaker endpoints, Sagemaker Hyperpod Serving), batching strategies, quantization, model parallelism, and latency/throughput tradeoffs
Provide deep technical guidance on training optimization distributed training strategies, framework selection (PyTorch, JAX, NeMo), SageMaker HyperPod, Slurm/PCS integration, checkpointing, and data pipeline design
Guide customers on GPU and accelerator profiling identifying bottlenecks (compute, memory, I/O), optimizing utilization, and tuning system-level performance
Help customers understand and apply model optimization techniques fine-tuning approaches (LoRA, QLoRA, full fine-tuning), RLHF/DPO, knowledge distillation, and efficient serving techniques (vLLM, TensorRT-LLM, Triton)
Responsibilities
- Work directly with the most important and exciting Startup customers in the GenAI model training and inference space, helping them adopt and scale large-scale workloads (e.g., frontier models, models, multi-modal systems, optimization) on AWS
- Advising customers on AI infrastructure requirements and trade-offs including GPU/Trainium selection, cluster topology, storage, networking (EFA), and cost optimization for training and inference
- Providing deep technical guidance on inference optimization model serving architectures (self-managed on EKS, SageMaker endpoints, Sagemaker Hyperpod Serving), batching strategies, quantization, model parallelism, and latency/throughput tradeoffs
- Providing deep technical guidance on training optimization distributed training strategies, framework selection (PyTorch, JAX, NeMo), SageMaker HyperPod, Slurm/PCS integration, checkpointing, and data pipeline design
- Guiding customers on GPU and accelerator profiling identifying bottlenecks (compute, memory, I/O), optimizing utilization, and tuning system-level performance
- Helping customers understand and apply model optimization techniques fine-tuning approaches (LoRA, QLoRA, full fine-tuning), RLHF/DPO, knowledge distillation, and efficient serving techniques (vLLM, TensorRT-LLM, Triton)
Requirements
Experience conveying complex technical concepts to both technical and business audiences
8+ years of experience in technology domain areas (e.g., systems engineering, cloud infrastructure, HPC, ML/AI, distributed computing)
3+ years of experience designing, implementing, or consulting on large-scale AI/ML infrastructure with hands-on experience on GPU-based computing, ML training infrastructure, and inference serving systems
Qualifications
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware, or experience with CUDA kernels or ML/low-level kernels
- Experience with vLLM, SGLang, TensorRT or similar platforms in production environments, or experience in performant kernel development (CUTLASS, FlashInfer)
- Experience with container orchestration for ML: EKS, Kubernetes operators for ML KubeRay, Karpenter, Keda, K8/DRA
- Experience with HPC schedulers and managed platforms: Slurm, AWS PCS (Parallel Computing Service), SageMaker HyperPod
- Experience with fine-tuning techniques: LoRA, QLoRA, RLHF, DPO, knowledge distillation, Quantization, KV optimization
Skills
Deep infrastructure and systems background combined with hands-on ML/AI expertise that enables you to lead engagements with Frontier AI labs, startups, and large enterprises
Understanding of the hardware layer: GPU architectures (NVIDIA A100/H100/B200, AWS Trainium/Inferentia), NVLink, EFA networking, storage hierarchies (FSx for Lustre, S3), and how they interact at scale
Orchestration layer: How to run large-scale training at least on one or more of EKS/Kubernetes, SageMaker HyperPod, Slurm/PCS — including cluster management, job scheduling, fault tolerance, and elastic scaling
Framework/model layer: Distributed training paradigms, inference frameworks (vLLM, llm-d, Triton, SGlang, etc), and optimization techniques (quantization, speculative decoding, KV-cache optimization)
Profiling and debugging layer: GPU profiling tools (NVIDIA Nsight, DCGM, PyTorch Profiler), identifying compute/memory/communication bottlenecks, and systematic performance tuning
Benefits
Comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave
Pay
Base salary range: $169,000.00 - $228,600.00 USD annually
Schedule
Full-time