Senior Inference Reliability Engineer
End-to-End Inference Reliability
Own the production health of customer inference workloads, including availability, request success, time to first token, inter-token latency, throughput, and operational efficiency.
Establish clear service-level indicators, objectives, performance baselines, and escalation paths for production endpoints.
Detection and Observability
Create telemetry, dashboards, alerts, and automated diagnostics needed to detect meaningful endpoint degradation before customers report it.
Create visibility across the full inference-serving path, including request queues, routing, scheduling, model servers, GPU utilization, networking, storage, and provider infrastructure.
Production Investigation
Lead the investigation of complex latency, throughput, capacity, and reliability regressions.
Determine whether an issue originates in customer traffic patterns, platform services, inference-engine configuration, GPU hardware, networking, storage, or an external infrastructure provider.
Remain accountable for the customer outcome while partnering with the appropriate engineering teams to implement the fix.
Incident Response and Prevention
Help lead customer-impacting incidents and establish effective operational practices for acknowledgement, diagnosis, recovery, and communication.
Convert significant incidents into automated tests, safeguards, runbooks, capacity controls, anomaly detection, and platform improvements.
Performance and Capacity
Partner with the LLM Performance team to validate that engine-level optimizations deliver measurable improvements in production.
Analyze workload behavior, capacity requirements, utilization, tail latency, and cost efficiency across heterogeneous GPU providers and hardware.
Help ensure that customer performance requirements are met without consuming unnecessary infrastructure capacity.
Production Feedback Loop
Identify recurring patterns across incidents, workloads, and customer escalations.
Translate those findings into improvements to the inference platform, reliability architecture, deployment processes, observability, and product roadmap.
Technical Abilities
Production Systems - Deep experience operating critical, customer-facing or business-critical production systems.
Ability to reason about service health across multiple layers rather than treating infrastructure availability as the complete customer outcome.
Reliability Engineering - Defining and operating service-level indicators and objectives, building actionable observability, leading incidents, performing failure analysis, and reducing mean time to detection and recovery.
Performance Diagnosis - Strong understanding of latency, throughput, queueing, resource contention, capacity, workload distribution, and tail-performance behavior.
Demonstrated ability to diagnose difficult production regressions and isolate bottlenecks across applications and infrastructure.
Distributed Systems - Knowledge of distributed-systems principles, including fault tolerance, scheduling, routing, load balancing, capacity management, consistency, and failure recovery.
Cloud and Infrastructure - Strong experience with Kubernetes, Linux, networking, storage, cloud infrastructure, and containerized production environments.
Experience operating across multiple cloud providers, regions, hardware configurations, or infrastructure suppliers is especially valuable.
Software Engineering - Ability to write production-quality software and build internal tooling, instrumentation, automation, and diagnostic systems.
Proficiency in languages such as Python, Go, Java, C++, or Rust.
Qualifications
5+ years of experience in production engineering, site reliability engineering, infrastructure engineering, distributed systems, ML infrastructure, database reliability, or performance engineering.
Demonstrated ownership of a critical production service or workload.
Experience diagnosing complex latency, throughput, capacity, or reliability problems across multiple system layers.
Strong software-engineering ability beyond infrastructure configuration and CI/CD automation.
Hands-on experience building observability, automation, diagnostic tooling, or production safeguards.
Strong communication and technical leadership skills, including the ability to coordinate incident resolution across engineering teams.
Experience with Kubernetes, Linux, networking, and cloud-native infrastructure.
While prior LLM-inference experience is not required, a demonstrated ability to learn unfamiliar systems and develop deep technical expertise is essential.