Jobs · Information Technology · California

Member of Technical Staff - Compute Platform

Reflection · San Francisco, CA · 1 mo ago
On-siteInformation TechnologyFull-time

About the role

The Compute Platform team at Reflection specializes in maintaining a K8s-based platform across multiple cloud environments. Your responsibilities will involve designing and iterating on the cluster management stack, implementing comprehensive monitoring and observability, and preparing the infrastructure for next-generation GPU deployments.

Responsibilities

  • Cluster Management: Build and maintain tools for automatic remediation, topology-aware scheduling, capacity planning, and rapid hardware debugging.
  • Platform Engineering: Design and refine the cluster management stack for large, multi-GPU fleets.
  • Misson: Implement comprehensive cluster-wide monitoring, focusing on durability and active performance benchmarking.
  • Roadmap Execution: Prepare the infrastructure for next-generation GPU deployments and increasingly larger cluster sizes.
  • Long-term: Own multi-cloud storage, petabyte-scale data replication, and GPU-to-GPU network performance.

Requirements

  • Systems-level engineering experience with a focus on cluster-wide behavior and maintenance.
  • Strong coding ability and a demonstrated focus on systems or GPU infrastructure.
  • Deep GPU hardware knowledge beyond standard Kubernetes, e.g., familiarity with NCCL.
  • Alignment with a K8s-first architecture.
  • Cloud storage expertise, specifically managing high-performance data products (like VAST) across multiple data centers, connecting those storage environments together and handling datasets and checkpointing at scale.

Qualifications

  • Systems-level engineering experience with a focus on cluster-wide behavior and maintenance.
  • Strong coding ability and a demonstrated focus on systems or GPU infrastructure.
  • Deep GPU hardware knowledge beyond standard Kubernetes, e.g., familiarity with NCCL.
  • Alignment with a K8s-first architecture.
  • Cloud storage expertise, specifically managing high-performance data products (like VAST) across multiple data centers, connecting those storage environments together and handling datasets and checkpointing at scale.

Skills

  • Systems-level engineering experience.
  • Strong coding ability.
  • Familiarity with GPU hardware and NCCL.
  • Experience with K8s and cloud storage.

Benefits

  • Top-tier compensation: Salary and equity structured to recognize and retain our talent globally.
  • Stock options: Everyone who joins and contributes to Reflection's success gets to share in the upside through stock options.
  • Health & wellness: Comprehensive medical, dental, vision, and life insurance, with an annual wellness allowance.
  • Meals: Lunch and dinner are provided in the office daily.
  • Life & family: 22 weeks paid parental leave for all new birthing and non-birthing parents, including adoptive and surrogate journeys.
  • Vacation days: Unlimited paid time off in the U.S. and 30 days in the U.K.
  • Sponsorship support: We sponsor visas to help exceptional talent join our team and support long-term immigration pathways where applicable.
  • Team building: We have regular off-sites, happy hours, and team celebrations.

Pay

Top-tier compensation: Salary and equity structured to recognize and retain our talent globally.

Schedule

Unlimited paid time off in the U.S. and 30 days in the U.K.

Similar jobs