AI Accelerator, Software Principal Engineer- Full-Stack
Ampere · Santa Clara, CA · 1 mo ago
HybridEngineering$195k–$292k/yrFull-time
About the role
We are looking for an engineer with strong experience in PyTorch-based AI deployment, accelerated inference execution, and systems integration across software components. The role involves working on temporal and multi-modal workloads, optimizing execution on target platforms, and building infrastructure to run AI models reliably in production environments.
Responsibilities
- Deploy and validate different AI models across supported inference environments (local, on-prem, or edge/accelerated platforms), optimizing runtime performance and reliability.
- Develop and tune inference graph transformations, including torch.export-based graph workflows.
- Collaborate with platform and infrastructure teams to improve model execution efficiency on supported AI accelerators and compute environments.
- Integrate model inference pipelines into scalable services, including batching, streaming, and runtime orchestration.
- Build and maintain middleware and communication layers to support modular, scalable system integration.
- Support long-term platform development for end-to-end inference services and production readiness.
Requirements
- Relevant Technical Areas:
- Temporal model architectures
- Multi-frame or sequence embeddings (e.g., video/text sequences)
- Attention-based models
- Multi-modal workloads (e.g., text, imaging, and other feature modalities)
- Automated labeling, evaluation, and validation workflows
- CPU/runtime performance optimization for system components
- Publish-subscribe middleware and distributed communication systems
Qualifications
- Bachelor's degree in Computer Science, Mathematics or a related technical field & 8 years of related experience; or Master's degree & 6 years
- Strong hands-on experience with PyTorch
- Experience deploying AI models to accelerated or constrained environments (e.g., edge, cloud GPU, or specialized accelerators)
- Familiarity with graph optimization and model-performance tuning
- Experience working with middleware, messaging, or distributed communication layers
- Good understanding of hardware/software interaction in AI systems
- Experience collaborating with hardware or platform partners
Skills
- Hands-on experience with PyTorch
- Experience deploying AI models to accelerated or constrained environments
- Familiarity with graph optimization and model-performance tuning
- Experience working with middleware, messaging, or distributed communication layers
- Good understanding of hardware/software interaction in AI systems
- Experience collaborating with hardware or platform partners
Benefits
- Premium medical insurance
- Dental insurance
- Vision insurance
- Income protection
- 401K retirement plan
Pay
- The full base pay range for this role is between $195,000 and $292,000.
Schedule
- Flexible work schedule
- 10+ paid holidays