Jobs · Engineering · California

AI Accelerator, Software Principal Engineer- Full-Stack

Ampere · Santa Clara, CA · 1 mo ago
HybridEngineering$195k–$292k/yrFull-time

About the role

We are looking for an engineer with strong experience in PyTorch-based AI deployment, accelerated inference execution, and systems integration across software components. The role involves working on temporal and multi-modal workloads, optimizing execution on target platforms, and building infrastructure to run AI models reliably in production environments.

Responsibilities

  • Deploy and validate different AI models across supported inference environments (local, on-prem, or edge/accelerated platforms), optimizing runtime performance and reliability.
  • Develop and tune inference graph transformations, including torch.export-based graph workflows.
  • Collaborate with platform and infrastructure teams to improve model execution efficiency on supported AI accelerators and compute environments.
  • Integrate model inference pipelines into scalable services, including batching, streaming, and runtime orchestration.
  • Build and maintain middleware and communication layers to support modular, scalable system integration.
  • Support long-term platform development for end-to-end inference services and production readiness.

Requirements

  • Relevant Technical Areas:
    • Temporal model architectures
    • Multi-frame or sequence embeddings (e.g., video/text sequences)
    • Attention-based models
    • Multi-modal workloads (e.g., text, imaging, and other feature modalities)
    • Automated labeling, evaluation, and validation workflows
    • CPU/runtime performance optimization for system components
    • Publish-subscribe middleware and distributed communication systems

Qualifications

  • Bachelor's degree in Computer Science, Mathematics or a related technical field & 8 years of related experience; or Master's degree & 6 years
  • Strong hands-on experience with PyTorch
  • Experience deploying AI models to accelerated or constrained environments (e.g., edge, cloud GPU, or specialized accelerators)
  • Familiarity with graph optimization and model-performance tuning
  • Experience working with middleware, messaging, or distributed communication layers
  • Good understanding of hardware/software interaction in AI systems
  • Experience collaborating with hardware or platform partners

Skills

  • Hands-on experience with PyTorch
  • Experience deploying AI models to accelerated or constrained environments
  • Familiarity with graph optimization and model-performance tuning
  • Experience working with middleware, messaging, or distributed communication layers
  • Good understanding of hardware/software interaction in AI systems
  • Experience collaborating with hardware or platform partners

Benefits

  • Premium medical insurance
  • Dental insurance
  • Vision insurance
  • Income protection
  • 401K retirement plan

Pay

  • The full base pay range for this role is between $195,000 and $292,000.

Schedule

  • Flexible work schedule
  • 10+ paid holidays

Similar jobs