Member of Technical Staff
Morph · San Francisco, CA · 1 mo ago
On-siteInformation TechnologyFull-time
About the role
Morph builds the inference infrastructure behind the fastest open models. Our stack spans kernels, model serving, routing, autoscaling, and capacity. We are hiring a performance engineer to make the entire system faster, cheaper, and more reliable.
Responsibilities
- Find the gap between theoretical hardware performance and production performance
- Trace latency and throughput regressions from the API layer down to individual kernels
- Optimize batching, scheduling, routing, quantization, and distributed execution
- Build benchmarks and observability that make bottlenecks obvious
- Validate that every optimization preserves model quality and correctness
- Stack-rank opportunities and ship the highest-impact fixes yourself
Qualifications
- Have optimized complex production systems
- Understand GPU performance, memory bandwidth, collectives, and inference serving
- Are strong in Python and comfortable navigating unfamiliar codebases
- Can turn profiling data into clear engineering decisions
- Care about tokens per second, tokens per dollar, and correctness equally
About the team
You will work directly with the founders on problems that determine how efficiently frontier-scale models can be served. Small team, enormous compute, immediate production impact.