Member of Technical Staff - Compute Platform
Reflection · New York, NY · 1 mo ago
On-siteEngineeringFull-time
About the role
The Compute Platform team at Reflection specializes in maintaining a K8s-based platform across multiple cloud environments. Your responsibilities will involve building and maintaining tools for automatic remediation, topology-aware scheduling, capacity planning, and rapid hardware debugging. You will also work closely with the training teams to design fault tolerance, node health checks, and remediation strategies.
Responsibilities
- Cluster Management: Develop and maintain tools for automatic remediation, topology-aware scheduling, capacity planning, and rapid hardware debugging.
- Platform Engineering: Design and iterate on the cluster management stack for large, multi-GPU fleets.
- Misson Monitoring & Observability: Implement comprehensive cluster-wide monitoring, focusing on durability and active performance benchmarking.
- Roadmap Execution: Prepare the infrastructure for next-generation GPU deployments and increasingly larger cluster sizes.
- Long-term Ownership: Help own multi-cloud storage, petabyte-scale data replication, and GPU-to-GPU network performance.
Requirements
- Systems-level engineering experience with a focus on cluster-wide behavior and maintenance.
- Strong coding ability and a demonstrated focus on systems or GPU infrastructure.
- Deep GPU hardware knowledge beyond standard Kubernetes, e.g., familiarity with NCCL.
- Alignment with a K8s-first architecture.
- Cloud storage expertise, specifically managing high-performance data products (like VAST) across multiple data centers, connecting those storage environments together and handling datasets and checkpointing at scale.
Qualifications
- Systems-level engineering experience with a focus on cluster-wide behavior and maintenance.
- Strong coding ability and a demonstrated focus on systems or GPU infrastructure.
- Deep GPU hardware knowledge beyond standard Kubernetes, e.g., familiarity with NCCL.
- Alignment with a K8s-first architecture.
- Cloud storage expertise, specifically managing high-performance data products (like VAST) across multiple data centers, connecting those storage environments together and handling datasets and checkpointing at scale.
Skills
- Systems-level engineering experience with a focus on cluster-wide behavior and maintenance.
- Strong coding ability and a demonstrated focus on systems or GPU infrastructure.
- Deep GPU hardware knowledge beyond standard Kubernetes, e.g., familiarity with NCCL.
- Alignment with a K8s-first architecture.
- Cloud storage expertise, specifically managing high-performance data products (like VAST) across multiple data centers, connecting those storage environments together and handling datasets and checkpointing at scale.
Benefits
- Top-tier compensation: Salary and equity structured to recognize and retain our talent globally.
- Stock options: Everyone who joins and contributes to Reflection's success gets to share in the upside through stock options.
- Health & wellness: Comprehensive medical, dental, vision, and life, with an annual wellness allowance.
- Meals: Lunch and dinner are provided in the office daily.
- Life & family: 22 weeks paid parental leave for all new birthing and non-birthing parents, including adoptive and surrogate journeys.
- Vacation days: Unlimited paid time off in the U.S. and 30 days in the U.K.
- Sponsorship support: We sponsor visas to help exceptional talent join our team and support long-term immigration pathways where applicable.
- Team building: We have regular off-sites, happy hours, and team celebrations.
Pay
Top-tier compensation: Salary and equity structured to recognize and retain our talent globally.
Schedule
Unlimited paid time off in the U.S. and 30 days in the U.K.