Jobs · Engineering · California

Senior Site Reliability Engineer - Data Infrastructure (San Jose)

ByteDance · San Jose, CA · 3 wk ago
On-siteEngineering$213k–$388k/yrFull-time

About the role

The Data Infrastructure SRE team is responsible for the reliability, scalability, and efficiency of the core data services that power our products. We manage a massive, distributed environment built on technologies like Kubernetes, Redis, MySQL, and Message Queue. Our work focuses on engineering the resilience and performance of the underlying platform that all product teams depend on. This role includes participation in a rotational on-call schedule to ensure nonstop coverage for our critical data infrastructure.

Responsibilities

  • Incident response and postmortems: Act as an incident commander for critical production issues, guiding the team through triage and resolution. Drive deep, blameless post-incident reviews and ensure follow-up actions are implemented to prevent recurrence.
  • SLO/SLA and error budgets: Define, negotiate, and maintain Service Level Objectives (SLOs) for critical data services. Champion the use of error budgets to balance reliability work with feature development.
  • Capacity and cost optimization: Lead initiatives in capacity planning, performance tuning, and resource management. Develop strategies and automation to ensure our infrastructure scales efficiently and stays within budget.
  • Pragmatic automation and AI orchestration: Design and build automation and leverage AI Agents to eliminate operational toil, improve deployment safety, and enhance overall operational efficiency. Focus on creating maintainable, robust tools and intelligent workflows that make the entire team more effective.
  • Operational excellence and change management: Uphold and improve our standards for production operations, including runbooks, monitoring, and alerting. Vet complex changes and deployments to ensure they meet our bar for production readiness.
  • Data Center and AI Infrastructure: Lead the construction, maintenance, and optimization of data centers and specialized AI infrastructure, ensuring high availability and peak performance for complex AI-driven workloads.
  • Cross-team influence and mentorship: Act as a subject matter expert on reliability, consulting with application development and other infrastructure teams. Mentor junior SREs, helping them develop their technical and operational skills.

Requirements

Minimum Qualifications:

  • Bachelor’s degree in Computer Science, a related technical field, or equivalent practical experience.
  • 5+ years of experience in a Site Reliability Engineering, Production Engineering, or similar role.
  • Strong proficiency in a programming or scripting language (e.g., Go, Python, Bash) for automation and tool development.
  • Deep understanding of Linux/Unix operating systems, networking fundamentals (TCP/IP, DNS), and distributed systems.

Preferred Qualifications:

  • Extensive hands-on experience managing large-scale data infrastructure (e.g., MySQL, Redis, Kafka, Flink).
  • Proven experience with container orchestration technologies, particularly Kubernetes, in a production environment.
  • Expertise in designing, analyzing, and troubleshooting large-scale distributed systems.
  • A systematic problem-solving approach, coupled with strong communication skills and a sense of ownership.
  • Experience leading incident response for complex, high-impact events.
  • Experience in the operation and construction of Data Centers is a plus.

About Us

Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut, and Pico, as well as platforms specific to the China market, such as Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.

Our innovative products are built to help people authentically express themselves, discover, and connect. We create value for our communities, inspire creativity, and enrich life. As ByteDancers, we strive to do great things with great people, leading with curiosity, humility, and a desire to make an impact in a rapidly growing tech company. By fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our company, and our users.

Benefits

  • Day one access to medical, dental, and vision insurance.
  • 401(k) savings plan with company match.
  • Paid parental leave, short-term and long-term disability coverage, and life insurance.
  • Wellbeing benefits.
  • 10 paid holidays per year.
  • 10 paid sick days per year.
  • 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).
  • Eligibility for additional discretionary bonuses/incentives and restricted stock units.

Pay

The base salary range for this position is $212,800 - $387,600 annually. Compensation may vary depending on qualifications, skills, competencies, experience, and location.

Similar jobs