Jobs · Engineering · California

Principal Software Engineer, Compute Fleet Management

OneClick Smart Resume · San Mateo, CA · Yesterday
Engineering$345k–$399k/yrFull-time

About the role

A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone.

Responsibilities

  • Serve as the overall technical lead for three Fleet Management pods, setting and aligning the technical direction across low-level provisioning, the data plane, and the control plane and product surfaces above them.
  • Architect the declarative, Kubernetes-style control planes that operate Roblox's compute fleet across on-prem and cloud, and define how capacity is provisioned, reconciled, and exposed at scale.
  • Own the design of the internal customer contracts and APIs that govern automation across the fleet, so that every infrastructure team can operate capacity safely and predictably.
  • Drive the strategy for self-serve capacity, including the internal-facing products and UIs that let teams request, manage, and reason about the compute they depend on.
  • Centralize and raise the bar on security, maintenance operations, and the uptime of all Roblox Kubernetes clusters, defining how fleet-wide changes ship reliably without impacting production.
  • Partner broadly with stakeholders inside and outside infrastructure to understand compute needs and drive innovation for our backend services, AI, and edge computing.
  • Write code daily, staying deep in the systems your org owns and leading by example on the hardest design and implementation problems.

Requirements

  • 10+ years of experience building and operating large-scale distributed systems and infrastructure.
  • A track record as the technical anchor an organization relies on, with the leadership to set direction across multiple teams and up-level the engineers around you.
  • Strong proficiency in Go, with deep experience designing and operating production services at fleet scale.
  • Hands-on experience building declarative, Kubernetes-style control planes and the reconciliation patterns behind them.
  • Strong proficiency with gRPC for service-to-service APIs and with SQL and Postgres for durable, high-scale state.
  • Experience operating compute capacity across both on-prem data centers and cloud providers, and a feel for the realities of running fleets at the scale of hundreds of thousands of instances.
  • A history of being highly cross-functional, partnering with stakeholders across and beyond infrastructure to design systems that keep compute supply and demand in balance.

Qualifications

  • Master’s degree in Computer Science, Engineering, or a related field.
  • Proven ability to mentor and develop technical teams.
  • Experience with cloud-native technologies and container orchestration systems like Kubernetes.
  • Strong understanding of cloud security principles and practices.
  • Excellent communication and collaboration skills.

Skills

  • Go programming language
  • Kubernetes
  • gRPC
  • SQL and Postgres
  • Cloud-native technologies
  • Container orchestration systems
  • Cloud security principles and practices

Benefits

  • Equity compensation
  • Comprehensive benefits package

Pay

The starting base pay for this position is $345,040—$399,420 USD. The actual base pay is dependent upon a variety of job-related factors such as professional background, training, work experience, location, business needs and market demand. Therefore, in some circumstances, the actual salary could fall outside of this expected range.

Schedule

Roles that are based in an office are onsite Tuesday, Wednesday, and Thursday, with optional presence on Monday and Friday (unless otherwise noted).

Similar jobs