Jobs · Management · New York

Production System Engineer Graduate (Server Management) - 2027 Start

ByteDance · New York, United States · 3 days ago
Management$76k–$128k/yrFull-time

Founded in 2012, ByteDance’s mission is to inspire creativity and enrich life. With a suite of products including TikTok, Lemon8, CapCut, and Pico, ByteDance makes it easier and more fun for people to connect with, consume, and create content.

About the role

The Server Management DevOps team is responsible for the end-to-end lifecycle management of servers across our self-built data centers in the United States and Europe. Our scope covers the complete server lifecycle, including new hardware introduction, data center delivery, production operations, hardware maintenance, configuration changes, capacity migration, asset decommissioning, data sanitization, and hardware reuse.

The team serves as a central coordination point between multiple functions, including:

  • Hardware New Product Introduction (NPI)
  • Server and data center operations
  • Field maintenance and infrastructure management
  • Hardware vendors and service providers
  • Supply chain and asset management
  • Infrastructure platform and automation engineering teams

Our goal is to ensure that server infrastructure operates reliably, efficiently, and compliantly at scale throughout its entire lifecycle.

Responsibilities

  • Server Infrastructure Operations: Assist with the deployment, validation, monitoring, maintenance, and lifecycle management of large-scale server fleets, including CPU and GPU servers.
  • Automation Development: Develop scripts, tools, and automation solutions using Python, Bash, Go, or other programming languages to reduce manual operational work and improve infrastructure efficiency.
  • Linux Systems: Work with Linux-based production environments and help troubleshoot operating system, hardware, storage, networking, and performance-related issues.
  • GPU and AI Infrastructure: Gain exposure to modern AI infrastructure and GPU server platforms, and contribute to operational tooling, validation, monitoring, or reliability improvements.
  • Monitoring and Data Analysis: Analyze server health, hardware failures, operational metrics, and infrastructure data to identify trends, risks, and opportunities for improvement.
  • AI for Infrastructure Operations: Explore opportunities to apply AI and large language models to infrastructure troubleshooting, automation, knowledge management, and operational decision-making.

Successful candidates must be able to commit to an onboarding date by the end of the year.

Requirements

  • Strong analytical and troubleshooting skills with the ability to learn unfamiliar technologies quickly.
  • Good communication skills and the ability to collaborate effectively in cross-functional engineering teams.

Qualifications

Minimum Qualifications

  • Individuals who are completing or have recently completed a Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, Information Technology, or a related technical field.
  • Experience in systems engineering, infrastructure operations, DevOps, Site Reliability Engineering, or related technical roles, or equivalent hands-on project experience.
  • Solid understanding of Linux system (Debian or Ubuntu preferred) administration and troubleshooting.
  • Programming or scripting experience in Python, Bash, Go, or another modern programming language.
  • Understanding of operating systems, computer architecture, networking fundamentals, and storage systems.

Preferred Qualifications

  • Experience developing automation tools or infrastructure software using Python, Bash, Go, or similar languages.
  • Working with server hardware, PC building, homelabs, or data center infrastructure.
  • Experience working with NVIDIA GPU platforms, AI infrastructure, CUDA, or high-performance computing environments.
  • Applying networking fundamentals, including TCP/IP, DNS, DHCP, VLANs, and routing.
  • Hands-on experience through internships, research, open-source projects, homelabs, technical competitions, or personal engineering projects.
  • Experience with infrastructure monitoring, observability, logging, or telemetry platforms.
  • Experience with Git, Docker, Kubernetes, REST APIs, SQL, or infrastructure automation frameworks such as Ansible.

Pay

The base salary range for this position is $76,000 - $128,000 annually. Compensation may vary depending on qualifications, skills, competencies, experience, and location. Base pay is one part of the total package that may include additional discretionary bonuses/incentives and restricted stock units.

Benefits

  • Day one access to medical, dental, and vision insurance.
  • 401(k) savings plan with company match.
  • Paid parental leave.
  • Short-term and long-term disability coverage.
  • Life insurance.
  • Wellbeing benefits.
  • 10 paid holidays per year.
  • 10 paid sick days per year.
  • 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

Similar jobs