Senior Network Reliability Engineer
About the role
The Platform Engineering Services team at Group 1001 is building a Site Reliability Engineering practice with a network scope. We're hiring a Sr. Network Reliability Engineer who will apply SRE principles — code-as-source-of-truth, SLOs and error budgets, alerting on symptoms rather than causes, failure-mode-first design, and the elimination of toil — to the firm's network platform from carrier edge through cloud fabric to Kubernetes pod boundary. This role focuses on systematically engineering the lights-on work out of existence, building abstractions that let other engineering teams express network intent in code, and treating the network as a single engineered system rather than a collection of vendor consoles. You will operate inside a DevSecOps practice spanning multi-cloud, multi-region environments, and partner closely with Cloud and Data Platforms, the NOC/SOC, and Cyber Security to extend reliability practice across the firm.
Responsibilities
- Treat reliability as an engineered property. Define SLOs and error budgets for the network platform — DNS resolution, edge availability, mesh ingress success, cross-region path health — and use them to gate changes, not just to color dashboards.
- Lead postmortems with a focus on permanent remediation, not pattern-recognition. Alert on symptoms users feel, not on causes that may or may not produce impact.
- Move network state into code. Use Terraform (or Pulumi), Ansible, and Python to replace CLI-driven configuration with declarative, version-controlled, peer-reviewed change running through Infra CI/CD. This applies equally to the edge tier (Cloudflare), security platforms (Zscaler ZIA/ZPA, ZTNA policies, next-gen firewalls), the cloud network fabric (Transit Gateway, Cloud WAN, VPCs, Route53, IPAM), and increasingly the Kubernetes and service-mesh layer.
- Build network policy as intent, not rule lists. Express what flows are permitted, what segments are isolated, what egress is inspected, what zones share DNS — and engineer the compilers that turn that intent into per-vendor configuration. Use Policy as Code (OPA/Rego, Sentinel, Cilium NetworkPolicy) to catch invariant violations at plan time, not apply time.
- Design, deploy, and manage network infrastructure using Infrastructure as Code (IaC) tools like Terraform or Ansible, moving the firm away from manual configuration to a code-first approach.
- Operate and extend the multi-account AWS Landing Zone — Cloud WAN segmentation, Transit Gateway peering, IPAM-driven CIDR allocation, shared private DNS, cross-account telemetry pipelines. Build platform abstractions that ensure new accounts or services land correctly with policy and connectivity composed from declarative inputs.
- Extend platform thinking into the container tier, covering Kubernetes networking, service mesh (Istio, Linkerd, Consul Connect), eBPF-based observability and policy (Cilium, Hubble), and integration points where mesh-level authz meets cloud-tier identity.
- Improve telemetry and observability with intent. Build alerts as structured payloads with runbook links, suspected blast radius, and dependency-aware suppression. Author both system-health dashboards for operators and end-user monitoring dashboards that reflect actual user experience using tools like Grafana, Elastic, and Open Telemetry.
- Mentor and grow the team by providing technical guidance to junior engineers, fostering a culture of learning, and sharing patterns across Platform Engineering.
- Handle hardware when required, providing maintenance and configuration support for routers, switches, and firewalls at data centers and offices, applying code-first practices where possible.
- Serve as an escalation point for network issues, troubleshooting with a focus on root cause analysis and permanent remediation, and maintaining a documentation-first mindset.
- Reduce toil and hand off cleanly by authoring runbooks and SOPs for the NOC, packaging routine work for L1/L2 handoff, and coordinating across Data Platforms, NOC/SOC, and Cyber Security to spread reliability practices.
Requirements
- Deep understanding of TCP/IP, BGP, OSPF, VPNs, and SD-WAN architecture.
- Proven experience with Terraform (state management, modules) and Ansible (playbooks, roles) — or similar — in a production environment.
- Proficiency in Python for automation and API interaction, or similar.
- Hands-on experience with Cloudflare, Zscaler, and/or enterprise firewalls.
- Experience configuring monitoring tools (e.g., Datadog, Prometheus, Grafana) to create meaningful alerts and dashboards.
Nice to Have
- Service mesh experience (Istio, Linkerd, Consul Connect, Cilium).
- eBPF-based observability (Hubble, Pixie).
- AWS Multi-account landing zone tooling experience (AFT, Control Tower, or equivalent).
- Policy as Code experience (OPA/Rego, Sentinel, Cilium NetworkPolicy).
Professional Attributes
- Documentation First: A strong belief that a job isn't done until the documentation is written.
- Toil Reduction: A mindset that actively seeks to automate repetitive tasks.
- Hybrid Capability: Willingness to handle physical hardware tasks when required while maintaining a software-centric engineering mindset.
Pay
The base pay for this position ranges from $135,000/year in our lowest geographic market up to $190,000/year in our highest geographic market. Pay is based on factors such as market location, job-related skills, and experience.
Benefits
- Comprehensive health, dental, and vision insurance plan options for employees (and their families) who meet eligibility guidelines and work 30 hours or more weekly.
- Basic and Supplemental Life Insurance.
- Short and Long-Term Disability.
- Immediate access to the Company’s Employee Assistance Program and wellness programs — no enrollment required.
- 401K plan with matching contributions by the Company.