Staff Network Reliability Engineer, Cloud Operations
Skylo has pioneered a standards-based approach to satellite connectivity, connecting smartphones and IoT devices directly to satellites without special hardware. Our direct-to-device service is live on millions of activated devices across five continents, covering over 72 million square kilometers, in partnership with leading satellite operators, mobile network operators, Tier-1 chipset makers, and OEMs. At the heart of this is Skylo's commercial NTN vRAN: a 3GPP standards-based, cloud-native platform bridging terrestrial and satellite networks.
About the role
As a Staff Network Reliability Engineer, Cloud Operations, you are the Cloud Infrastructure domain authority within Skylo's production NTN network. You ensure the health of everything running on the infrastructure—RAN NFs, Core NFs, OSS, BSS, and the observability pipeline. You operate across Skylo's full hybrid cloud estate: GCP public cloud (GKE clusters, Pub/Sub pipelines, Cloud SQL) and on-premise private cloud infrastructure (bare-metal Kubernetes, hyperconverged compute, software-defined storage). At the Staff NRE level, you define and maintain SLOs tied to network SLA commitments, own error budget tracking, drive toil reduction, and partner with Ops Platform Engineering to automate infrastructure remediation.
Responsibilities
- Cloud Infrastructure Operations & Health Ownership
- Own 24x7 cloud infrastructure health across Skylo's hybrid production environment: GKE cluster node status, namespace and pod health, Persistent Volume Claim availability, network policies, and multi-cluster federation.
- Maintain on-premise Kubernetes cluster health: bare-metal node availability, container runtime stability, CNI networking, persistent storage arrays (Ceph/Rook), and hyperconverged compute platform operations (Harvester, KubeVirt, or KVM).
- Monitor and triage infrastructure alarms using OSS dashboards, Grafana/VictoriaMetrics telemetry, GCP Cloud Monitoring, and Loki log correlation.
- Execute and own Cloud Infra runbooks for P2–P4 fault categories: GKE node recovery, pod eviction and rescheduling, PVC repair, database failover execution, Prometheus WAL corruption recovery, ArgoCD drift remediation, and certificate rotation.
- Maintain BSS-IIS GKE cluster monitoring and infrastructure health; update runbooks to reflect current cluster topology after every infrastructure change.
- Observability Pipeline & Data Platform Operations
- Own the observability pipeline end-to-end: Prometheus scrape target integrity, VictoriaMetrics retention and query performance, Grafana dashboard coverage, OpenTelemetry collector health, and alert routing via Pub/Sub.
- Maintain database reliability: PostgreSQL streaming replication health, backup and restore procedures, failover testing, query performance monitoring; Redis cluster operations, eviction policy management, and persistence configuration.
- Ensure log aggregation pipeline health (Loki or ELK): ingestion rates, retention policies, query performance, and completeness.
- Partner with Network Implementation (NI) on all planned infrastructure changes: validate post-deployment observability and sign off on operational readiness.
- Cloud Infrastructure Incident Diagnosis & Escalation Authority
- Serve as the L3 escalation authority for all Cloud Infra incidents: diagnose at the Kubernetes, storage, network, and database layer using kubectl, GCP console, node logs, and infrastructure telemetry.
- Lead Cloud Infra troubleshooting bridges for GKE node failures, cluster upgrade failures, storage outages, PubSub pipeline disruptions, database failover events, and ArgoCD sync failures.
- Diagnose and resolve infrastructure failure modes: node NotReady conditions, pod CrashLoopBackOff chains, PVC mount failures, CSI driver errors, network policy misconfigurations, Helm release drift, etcd latency spikes, and cross-cluster federation breaks.
- Participate in the global 24x7 on-call rotation as the Cloud Infra domain escalation tier, functioning as the technical decision-maker for Sev 1 events.
- SLO Engineering & Reliability
- Define and maintain SLOs for all Cloud Infra components: GKE control plane availability, database query latency, storage IOPS, message pipeline throughput, and observability stack uptime—tied directly to network SLA commitments.
- Own error budget tracking and the process for trading error budget against deployment velocity; escalate when error budget burn rate requires engineering intervention or deployment freezes.
- Drive toil reduction: identify and eliminate manual Cloud Infra procedures; own the roadmap to automated cluster recovery, rolling restarts, storage repair, and certificate rotation.
- Lead capacity planning for compute, storage, and network resources across public and private cloud based on subscriber projections and new MNO partner onboarding.
- Root Cause Analysis & Post-Incident Ownership
- Own Cloud Infra RCA end-to-end: lead the investigation, document the complete causal chain, and deliver systemic action items with owners, timelines, and measurable success criteria.
- Deliver Initial RCA documentation within defined SLA windows; identify systemic infrastructure failure patterns and translate them into engineering requirements.
- Contribute to the weekly and monthly Network Performance Report: infrastructure availability, database latency trends, storage IOPS, observability pipeline health, and SLA deviation analysis.
- Runbook Authorship & Operational Standards
- Author, own, and maintain all Cloud Infra runbooks and SOPs: GKE node recovery, database failover, storage expansion, Prometheus WAL repair, ArgoCD rollback, certificate rotation, and cluster upgrade procedures.
- Define the diagnostic decision tree for each known infrastructure fault class: entry condition, triage steps, isolation method, resolution action, and escalation criteria.
- Validate and sign off on operational readiness for all infrastructure changes: DCI/DCE build-outs, GKE cluster expansions, Kubernetes version upgrades, and new on-premise hardware deployments.
- Cross-Functional Collaboration & Team Development
- Partner with Ops Platform Engineering (OPE) as the Cloud Infra domain's primary automation consumer: define Kubernetes event schemas, alert-to-action contracts, and closed-loop policy requirements.
- Represent Cloud Infra in NI architecture reviews: define observability and operational readiness requirements for all infrastructure expansions and GitOps pipeline changes.
- Collaborate with Core NRE and RAN NRE on infrastructure-layer issues affecting network functions: pod scheduling, PVC availability, network policy changes, and platform upgrade impacts.
- Partner with security teams on infrastructure hardening: patch compliance, RBAC policies, network segmentation, container image scanning, and runtime security monitoring.
- Surface toil and automation opportunities to the Service Assurance & Automation team; document procedures, frequency, and MTTR cost as structured input to the automation backlog.
- Mentor Senior NREs in Cloud Infra domain depth: Kubernetes troubleshooting patterns, storage operations, database reliability, observability pipeline internals, and escalation judgment.
- GitOps, IaC & Platform Engineering Interface
- Own operational oversight of GitOps tooling in production: ArgoCD sync health, Helm chart version management, drift detection, and rollback execution for multi-cluster deployments.
- Review and validate Infrastructure as Code (Terraform, Ansible) changes that impact production—ensure operational impact is assessed and observability is in place before merge.
- Engage Skylo's Platform Engineering and NI teams with full operational context when issues exceed operational resolution authority.
Requirements
- 8–10+ years of infrastructure engineering, Site Reliability Engineering, or cloud operations in a production 24x7 environment—with direct on-call ownership for Kubernetes-at-scale environments.
- Deep Kubernetes expertise: multi-cluster operations (GKE or EKS), node pool management, RBAC, network policies, persistent storage (PVC, CSI drivers), CRD/operator patterns, and production cluster upgrade procedures.
- Hybrid cloud operations: hands-on experience operating both public cloud (GCP or AWS) and on-premise/private cloud infrastructure (bare-metal Kubernetes, KVM, or hyperconverged platforms).
- Production observability stack ownership: Prometheus, VictoriaMetrics, Grafana, OpenTelemetry, and log aggregation (Loki or ELK).
- Database reliability engineering: PostgreSQL streaming replication, failover testing, query performance tuning, and Redis cluster operations.
- Infrastructure as Code (IaC) and GitOps: Terraform, Ansible, ArgoCD, Helm, and CI/CD pipelines for infrastructure changes.
- Scripting and automation: Python, Go, or Bash for infrastructure automation and tooling.
- Strong incident management and troubleshooting skills: experience leading complex incident bridges, diagnosing multi-layer failures, and driving to resolution.
- SLO/SLI engineering: experience defining and maintaining SLOs, error budget tracking, and toil reduction initiatives.
- Collaborative mindset: ability to work across engineering, security, and operations teams to drive operational excellence.