NOC Technician (Data Center and Site Ops)
About The Role
As a NOC Technician, you are the eyes and the voice of the site — never the hands. You staff the Network / Campus Operations Center and continuously observe site health signals across xAI campuses. You detect and verify campus-impacting events, assemble the right responders, run incident communications leadership can trust, and drive every major incident to a completed report and a tracked corrective project. You work with Site Reliability Engineering, SiteOps, Facilities, Hardware Failure Analysis, SWE Platforms, and vendors — escalating correctly the first time and maintaining the institutional memory across shifts and sites. One sentence: watch the campus, run the bridge, leave the wrench work and deep root cause to the teams that own them.
Responsibilities
- Continuous monitoring (the watch)
- Staff the console per shift schedule to sustain 24/7 coverage (coverage posture: 2 on console per site)
- Watch the designated signal surface: cluster health dashboards, node availability, network health, facility trend panels (power/cooling), storage alarms, and threshold breaches as defined by SRE monitoring standards
- Acknowledge every page/alert within the SLA; classify it (actionable / known / noise) and log the disposition; feed noise patterns back to SRE for suppression or redesign
- Maintain a live picture of ongoing maintenance, planned work, and degraded-but-accepted states so real anomalies stand out
- Detection, triage & escalation
- Detect → verify → escalate within defined time budgets; verification is signal-level (is it real, what's the blast radius), not deep diagnosis
- Operate the escalation matrix: NOC → on-call SRE → domain owners (SiteOps, Facilities, Network, Storage, HW FA, vendors); page correctly the first time
- Recommend incident declaration and severity to the on-call SRE; declare directly per runbook when thresholds are unambiguous
- Incident communications & coordination
- Open and run the bridge; get the right people on within the time-to-bridge SLA
- Own stakeholder communications: first update within the SLA, then a fixed cadence until resolution
- Maintain the incident timeline in real time — timestamps, actions, decisions, engagements
- Track who owns what during the incident and call out stalls
- First-pass RCA framing & closure
- Produce initial framing for major site outages: what happened, when it started, what's impacted (halls/racks/services), what changed recently, who is engaged
- Hand framing to SRE / Hardware FA for depth — the NOC does not publish root cause
- Write major-incident reports; open corrective projects in Linear with named owners and track them to closure ("filed" is not "done")
- Shift operations, runbooks & improvement
- Run structured shift handoffs and keep durable shift logs; maintain cross-site awareness
- Own and continuously improve NOC runbooks: escalation matrix, comms templates, severity ladders, per-signal response procedures
- Participate in game days run by SRE; every incident where the runbook was wrong or missing produces a runbook change before the incident closes
- Explicitly not this role
- Wrench work: swaps, reseats, physical recovery (SiteOps)
- Power / cooling / building plant operation (Facilities)
- Deep hardware root-cause analysis or vendor CAPA (Hardware Failure Analysis)
- Monitoring architecture, alert design, or technical SEV command (Site SRE)
- Building or operating reliability tooling such as SRT, turnback, or dashboards (SWE Platforms)
Basic Qualifications
- High school diploma or equivalency certificate
- 1+ year of professional experience in a Network Operations Center (NOC), Security Operations Center (SOC), mission-control / dispatch, data center operations watch, or equivalent 24/7 monitoring and incident-communications role
- Demonstrated written and verbal communication skills under time pressure (stakeholder updates, handoffs, timelines)
Preferred Skills And Experience
- Calm under pressure; excellent written and verbal communications — leadership should be able to trust your incident updates verbatim
- Pattern recognition across domains; multi-domain curiosity (compute, network, storage, power/cooling signals)
- Experience following and improving process: runbooks, escalation matrices, shift handoffs, post-incident follow-through
- Prior NOC, SOC, or critical-environment operations experience in a datacenter or hyperscale infrastructure environment
- Familiarity with reading operational dashboards, acknowledging/classifying alerts, and coordinating across on-site technicians, facilities, and engineering on-call
- Comfort with ticketing / project tracking systems (e.g. Linear, Jira) for opening and chasing corrective work to closure
- Industry certifications a plus (Network+, Security+, ITIL, or similar) — not a substitute for judgment and communications quality
- Basic familiarity with datacenter topology (racks, fabric, OOB) and how facility events affect compute availability — enough to triage and escalate correctly, not to deep-diagnose
- Bachelor's degree in IT, Computer Science, Cybersecurity, or STEM discipline preferred but not required
Additional Requirements
- Must be available for on-shift rotations supporting 24/7/365 console coverage
- Shift structure (e.g. 12-hour rotations) to be confirmed; nights, weekends, and holidays are part of the role
- Must be able to work extended hours during major incidents as needed
Success Looks Like
- Coverage attainment / shift fill rate vs plan
- Time-to-bridge for major incidents; first-update and cadence SLA attainment
- Page accuracy / escalation correctness
- % of major incidents with a complete timeline and follow-up projects tracked to done
- Not measured by: raw page counts, or heroics without a paired prevention item