Senior Network Reliability Engineer, Incident Management
Skylo · Mountain View, CA · 1 mo ago
HybridEngineeringFull-time
About the role
Skylo has pioneered a standards-based approach to satellite connectivity. Our direct-to-device service is live on millions of activated devices across five continents, covering more than 72 million square kilometers, in partnership with leading satellite operators, mobile network operators, Tier-1 chipset makers, and OEMs worldwide. We are looking for a Network Reliability Engineer (NRE) for Incident Management to join our Global Product Support & Customer Success organization.
Responsibilities
- Serve as the central command point during all network degradations, service disruptions, and subscriber-impacting events, opening the bridge, initiating the incident management process, and maintaining command from first alert through full-service restoration.
- Own end-to-end incident lifecycle management for Sev 1-4: initiate bridge calls, identify the cause and impacted domain, page the correct on-call NRE, maintain bridge discipline with clear ownership and timelines, and drive to restoration.
- Prioritize incidents according to urgency and business impact, classifying severity accurately using alarm signatures, subscriber impact data, and domain KPI telemetry from available OSS systems.
- Escalate to subject matter experts in Operations and Engineering teams when critical or time-sensitive resolution is required, providing full technical context, a structured problem statement, and a documented timeline.
- Engage and interface with vendor support teams (RAN vendor, Core vendor, cloud infrastructure) when incident resolution requires external escalation; track vendor SLA response and escalate vendor delays to the domain NRE.
- Support hypercare operations during major network launches, high-risk change windows, and special events, maintaining readiness and acting as first responder for any degradation during the hypercare window.
Qualifications
- 5-10+ years of experience in telecom/wireless operations, network operations, or NRE in a production 24x7 environment.
- Demonstrated ability to independently manage incident bridge calls: open the war room, maintain bridge discipline, drive to resolution, and produce a structured incident record.
- Strong understanding of telecom network environments and 5G functional components, with sufficient knowledge of RAN (CUSM, eCPRI, PTP/SyncE), 5G Core (AMF, SMF, UPF), and Cloud/OSS to triage intelligently and escalate with context.
- Incident and outage management expertise: ability to prioritize by urgency and impact, manage multiple simultaneous events, and operate effectively under high-pressure 24x7 conditions.
- Hands-on experience with at least one observability platform: Grafana dashboards, Prometheus alerting, Loki log queries, or equivalent, with the ability to independently navigate to relevant signals during an active incident.
- Kubernetes operational literacy: able to run kubectl get pods, describe a failing pod, read container logs, and identify health issues at the level needed to triage and escalate a platform-layer incident.
- Structured written communication: capable of producing clear incident timelines, executive stakeholder updates, and post-incident summaries under time pressure.
- Ticketing system proficiency (Jira, ServiceNow, or equivalent): incident lifecycle management, escalation workflows, and backlog hygiene.
- On-call tooling experience (PagerDuty or equivalent): alert acknowledgement, escalation policy management, and on-call scheduling.
- Ability to span departments and build strong working relationships with RAN, Core, Cloud, Engineering, and external partner teams to drive joint incident resolution.