Jobs · Information Technology · California

Staff Infrastructure Software Engineer, Enterprise AI

Scale AI · San Francisco Bay Area · 2 wk ago
HybridInformation Technology$216k–$270k/yrFull-time

Scale GP is building the infrastructure that makes enterprise AI seamless. We are looking for a Senior or Staff Infrastructure Engineer to act as a primary technical lead, engineering the "paved road" for our knowledge retrieval and inference engines. You won’t just be managing resources; you’ll be defining the deployment standards for Agentic workflows at scale. Your mission is to bridge the gap between complex AI orchestration and world-class infrastructure, ensuring our platform remains the most reliable destination for enterprise agents.

About the role

The ideal candidate thrives in a fast-paced environment, has a passion for both deep technical work and mentoring, and is capable of setting a long-term technical strategy for a critical domain while maintaining a strong, hands-on delivery focus. You will architect and implement solutions across multiple cloud providers (GCP, Azure, AWS) for customers in diverse, highly-regulated industries like healthcare, telecom, finance, and retail.

Responsibilities

  • Architect multi-cloud systems and abstractions to allow the SGP platform to run on top of existing Cloud providers.
  • Use our own data and AI platform to analyze build and test logs and metrics to identify areas for improvement.
  • Define the architectural patterns for our multi-cloud infrastructure to support secure, reliable, and scalable Agentic workflows for enterprise customers.
  • Enhance engineering and infrastructure efficiency, reliability, accuracy, and response times, including CI/CD processes, test frameworks, data quality assurance, end-to-end reconciliation, and anomaly detection.
  • Collaborate with platform and product teams to develop and implement innovative infrastructure that scales to meet evolving needs.
  • Design and champion highly scalable, reliable, and low-latency infrastructure and frameworks for building, orchestrating, and evaluating multi-agent systems at enterprise scale.
  • Lead the infrastructure roadmap with a strong focus on compliance, privacy, and security standards, including designing change management and data isolation strategies.
  • Own the development and maintenance of our best-in-class Agentic observability platform (logging, metrics, tracing, and analytics) to proactively ensure system health and enable rapid incident response.
  • Drive developer efficiency by building automated tooling and championing Infrastructure-as-Code (IaC) paradigms throughout the engineering organization to improve workflows and operational efficiency.

Requirements

  • Proven experience in a senior role, with 5+ years of full-time software engineering experience.
  • Deep understanding of modern infrastructure practices, including CI/CD, IaC (e.g., Terraform, Helm Charts), container orchestration (e.g., Kubernetes) and observability platforms (e.g., Datadog, Prometheus, Grafana).
  • Extensive experience with at least one major cloud provider (AWS, Azure, or GCP).
  • Strong knowledge of security and compliance in enterprise environments, with a focus on access management, data isolation, and customer-specific VPC setups.
  • Proficiency in Python or JavaScript/TypeScript, and SQL.

Qualifications

  • Bonus points: Hands-on experience and a passion for working with Agents, LLMs, vector databases, and other emerging AI technologies.

Benefits

  • Comprehensive health, dental and vision coverage.
  • Retirement benefits.
  • A learning and development stipend.
  • Generous PTO.
  • This role may be eligible for additional benefits such as a commuter stipend.

Pay

Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The base salary range for this full-time position in the locations of San Francisco, New York, and Seattle is: $216,200—$270,250 USD.

Similar jobs