Jobs · Engineering

Cloud Platform Engineer

SambaNova · San Jose, CA · 2 wk ago
RemoteRemoteEngineeringFull-time

About the role

As a Cloud Platform Engineer, you will specialize in our AI Inferencing Service and act as the guardian of its reliability, performance, and scalability. You will bridge the gap between software development and operations, applying an engineering mindset to solve operational challenges. Your primary focus will be ensuring our inference endpoints have exceptional uptime, low-latency response times, and efficient resource utilization, directly impacting the experience of our customers and the success of our AI products. This role includes participating in a shared on-call rotation to maintain 24/7 service reliability.

Responsibilities

  • Service Ownership & On-Call: Take shared ownership of the production inferencing service, including its availability, latency, performance, efficiency, change management, monitoring, emergency response, and capacity planning across multiple regions (e.g., Asia, Europe, Latin America). Implement and support AI infrastructure in new regions to support business growth.
  • On-Call & Work-Life Balance: Participate in a balanced on-call rotation (primary/secondary, follow-the-sun model) to provide 24/7 support. Focus on prevention through automation, robust testing, and system design to minimize incidents. Alerts are actionable and require immediate human intervention.
  • Incident Management: Lead responses to incidents affecting the inferencing service, drive blameless post-mortems, and implement corrective actions to prevent recurrence.
  • Monitoring & Alerting: Develop and maintain advanced monitoring, alerting, and dashboarding (using tools like Prometheus, Grafana, Datadog) to gain insights into service health, model performance (latency, throughput, error rates), and accelerator utilization. Ensure alerts are actionable with a low false-positive rate.
  • Performance & Scalability: Proactively identify and eliminate performance bottlenecks. Design and implement auto-scaling policies to handle variable inference loads cost-effectively. Use insights from incidents to enhance system stability and scalability.
  • Infrastructure as Code (IaC): Manage and evolve cloud infrastructure (AWS, GCP, Azure, and on-prem) using tools like Terraform and Ansible, ensuring it is secure, repeatable, and scalable.
  • CI/CD & Automation: Build and improve CI/CD pipelines for seamless and safe deployment of new model versions and service updates. Automate manual toil identified during on-call shifts to reduce operational overhead.
  • Capacity Planning: Forecast infrastructure needs based on product roadmaps and usage trends. Work with finance and engineering teams to manage cloud costs and optimize spending.
  • SLOs & SLIs: Define, measure, and report on Service Level Objectives (SLOs) and Indicators (SLIs) for the inferencing platform, using data to drive prioritization and reliability investments.

Requirements

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 3-5+ years of experience in a Site Reliability Engineer, DevOps, or related role supporting a large-scale, customer-facing service in a public cloud environment (AWS, GCP, Azure).
  • Strong programming/scripting skills in languages like Python, Go, or Java.
  • Proven experience with containerization and orchestration technologies (Docker, Kubernetes).
  • Deep understanding of monitoring and observability principles and tools (e.g., Prometheus, Grafana, ELK Stack, Datadog).
  • Solid experience with Infrastructure as Code (e.g., Terraform, CloudFormation).
  • Familiarity with CI/CD principles and tools (e.g., Jenkins, GitHub Actions, ArgoCD).
  • Excellent problem-solving skills and a systematic approach to troubleshooting complex distributed systems.

Skills

  • Nice-to-Haves:
    • Experience in a hybrid environment bridging cloud and on-premise/data center infrastructure.
    • Direct experience supporting ML/AI inferencing services in production.
    • Familiarity with GPU-accelerated computing and optimizing workloads for NVIDIA GPUs (mapping to RDUs).
    • Knowledge of model serving frameworks like vLLM, SGLang, or Ray.
    • Understanding of MLOps principles and practices.
    • Experience with managing and tuning databases (SQL or NoSQL) and caching systems (Redis, Memcached).
    • Strong Linux/Unix system administration fundamentals.

Benefits

  • Health & Wellness: 95% premium coverage for employee medical insurance, 77% for dependents, and a Health Savings Account (HSA) with employer contribution. Dental, Vision, Short/Long-term Disability, Basic Life, Voluntary Life, and AD&D insurance. Flexible Spending Account (FSA) options (Health Care, Limited Purpose, Dependent Care).
  • Well-being: Full subscription to Headspace, Gympass+ membership (access to physical gyms), One Medical membership, counseling services via Employee Assistance Program, and more.
  • Compensation: Competitive total rewards package including base salary, equity, and excellent benefits.
  • Work Environment: Flexible work environment with autonomy and opportunities for growth.

About SambaNova Systems

Join the company building the future of AI computing. At SambaNova, we disrupt the AI and high-performance computing space with our integrated hardware and software platform. Our DataScale systems and SambaFlow software push the boundaries of generative AI and large language models. We are a team of passionate innovators tackling some of the world's most challenging computational problems.

SambaNova Suite™ is the first full-stack, generative AI platform—from chip to model—optimized for enterprise and government organizations. Powered by the intelligent SN40L chip, the platform is fully integrated, delivered on-premises or in the cloud, and combined with state-of-the-art open-source models fine-tuned securely using customer data. Customers retain model ownership in perpetuity, turning generative AI into one of their most valuable assets.

Similar jobs

Cloud Platform Engineer

Hansen Talent Group (HTG)Columbia, SC· 2 wk ago
Information Technologyapply on www2.jobdiva.com