Sr. Principal Software Engineer
Ladders · United States · 2 wk ago
RemoteRemoteEngineering$141k–$226k/yrFull-time
For a leader in the Manufacturing & Automotive space, this role leads work at the intersection of data, AI-enabled capabilities, and scalable technology delivery. You will collaborate across engineering, product, operations, and business stakeholders to translate complex requirements into practical technology solutions, influencing architecture, execution quality, and long-term technology capabilities in a manufacturing environment.
Location: Remote – US-based candidates only, no visa sponsorship available.
Responsibilities
- Optimize and deploy LLM inference pipelines
- Manage inference runtimes across various platforms
- Enhance model performance using techniques like quantization and kernel fusion
- Drive improvements in latency and throughput for production products
- Enable efficient deployment independent of external vendors
- Build expertise in key inference engines
- Adapt runtimes for constrained environments
Qualifications
- 5-7 years experience in ML inference performance optimization
- In-depth knowledge of GPU architecture and memory hierarchies
- Proficiency in CUDA and low-level performance tuning
- Experience deploying models in production environments
- Familiarity with inference engines like vLLM, TensorRT-LLM, llama.cpp, QAIRT
- Expertise in quantization techniques including INT8, INT4, FP4, FP8, AWQ, GPTQ
- A solid understanding of latency optimization strategies
Pay
$141,400 – $226,300 annually
Benefits
- Annual bonus opportunity
- Comprehensive insurance coverage including medical, dental, vision, life, and disability
- Paid time off and holidays
- Company contributions to RRSP
- Equity awards available for select roles
- Remote or hybrid work options based on position