Jobs · Information Technology · Alabama

Data Center Production Operations Engineer

Meta · Huntsville, AL · 2 wk ago
Information Technology$84k–$130k/yrFull-time

About the role

Meta is seeking a forward-thinking, experienced engineer to join the Production Operations team within our Data Centers. These Data Centers are the foundation upon which our rapidly scaling infrastructure efficiently operates and delivers innovative services. Meta leads the global data center industry in design and operations. This role thrives in a fast-paced, technical environment where adaptability and flexibility are key to success. We seek an IT professional with advanced, hands-on technical skills in server hardware and Linux, ideally in a Data Center environment. Broad knowledge of server administration and experience in large-scale distributed data center projects are core competencies. The candidate should also have working knowledge in a few of these areas: Hardware repair, OS management, Tooling and Automation, Networking, or Technical Project Management.

Responsibilities

  • Support platform health by resolving and closing tickets, addressing root causes through remote troubleshooting and physical inspection of services in data halls.
  • Participate in deep dives and root cause analysis of highly technical issues, including automated tooling, hardware failures, and network issues.
  • Collaborate with cross-functional teams on projects related to process, hardware, and automation.
  • Serve as the point of contact for introducing new platforms and hardware, accelerating their transition to sustained mass production.
  • Use tools and data analysis to identify issues, communicate with stakeholders, and escalate as needed.
  • Identify corrective actions for hardware issues, work with internal teams and vendors, and influence future design changes for serviceability.
  • Solve systemic hardware and software issues at scale using scripting, automation, and tooling for global resolution.
  • Continuously evaluate and improve processes, tools, and systems to optimize efficiency and repair quality.
  • Use data analytics to maximize server uptime and utilization, understanding hardware failure rates and service level agreements.
  • Support and train team members to identify better ways to resolve issues and update tools and processes.
  • Provide engineering support and serve as a technical resource for the team, leadership, and cross-functional partners.
  • Maintain and update documentation, including procedures, runbooks, and guides.
  • Build cross-functional relationships and influence policies to improve global data center operations.
  • Participate in a 24/7 on-call rotation.
  • Travel up to 15% of the time.

Requirements

  • BS, BA, or BEng in a technical field or commensurate experience.
  • 5+ years of technical IT experience in an infrastructure environment, such as Systems Administrator, DevOps Engineer, or Site Reliability Engineer.
  • Intermediate-level understanding of Linux (or equivalent OS) in a complex IT environment, with the ability to triage, debug, and troubleshoot server issues.
  • Hands-on experience with server hardware and components, including storage.
  • Intermediate-level knowledge of data center functions and technologies, including electrical, cooling, structured cabling, security, and network.
  • Experience managing technical issues and driving root cause analysis.
  • Experience participating in technical projects related to process improvement, technology, or automation.
  • Ability to communicate effectively and tailor messages to the audience.
  • Intermediate-level knowledge of technologies such as HTTP, DNS, RAID, and DHCP.
  • Experience providing technical guidance to external vendors.
  • Experience debugging, modifying, or developing scripts in at least one of these languages: Bash, PHP, Python, SQL, Rust, Go, or Perl.
  • Knowledge of out-of-band/lights-out server communication methods, such as IPMI and serial console.
  • Experience using data and metrics to drive decisions.

Preferred Qualifications

  • Experience fostering growth in others and driving influence across organizational levels.
  • Experience in a large-scale data center environment.
  • Experience with large-scale AI implementations.
  • Six Sigma knowledge or certification.

About Meta

Meta builds technologies that help people connect, find communities, and grow businesses. Since launching Facebook in 2004, we’ve expanded our apps to include Messenger, Instagram, and WhatsApp, empowering billions worldwide. Now, Meta is moving beyond 2D screens into immersive experiences like augmented and virtual reality to shape the next evolution in social technology. By building with us, you help create a future beyond digital constraints—transcending screens, distance, and even physics.

Pay

$83,990/year to $130,000/year + bonus + equity + benefits. Individual compensation is determined by skills, qualifications, experience, and location.

Similar jobs