Senior Principal Network Engineer
About Us
Graphcore is one of the world’s leading innovators in Artificial Intelligence compute, developing hardware, software, and systems infrastructure to unlock the next generation of AI breakthroughs and power widespread adoption across industries. As part of the SoftBank Group, Graphcore joins an elite family of companies driving transformative technologies with a shared vision: enabling Artificial Super Intelligence and ensuring its benefits are accessible to everyone.
Graphcore’s teams bring diverse backgrounds and skills—AI research specialists, silicon designers, software engineers, and systems architects—fostering a culture of continuous learning and innovation.
The Team
The Data Center Network Engineering team designs and operates the high-performance network fabrics powering Graphcore’s AI compute platforms. Collaborating with hardware engineering, AI researchers, and infrastructure teams, the team builds scalable networking environments optimized for distributed training and inference workloads. Key focus areas include high-speed Ethernet fabrics, lossless networking, RDMA transport, and large-scale automation frameworks for next-generation AI clusters.
Responsibilities
- Assist in defining ultra-high-bandwidth, non-blocking AI network fabrics (Clos spine-leaf-super-spine architectures) for large-scale distributed AI workloads.
- Optimize performance of lossless Ethernet fabrics using congestion control mechanisms such as PFC, ECN, and DCQCN to support RDMA/RoCEv2 communication.
- Lead initiatives to implement NetDevOps practices and develop automation for provisioning, configuration management, and network remediation.
- Design and deploy high-resolution telemetry pipelines to monitor network health, detect microbursts, and analyze congestion patterns.
- Support modeling, deployment, configuration, and monitoring of data center network fabrics including scale-out, scale-up, and front-end networks.
- Collaborate cross-functionally with hardware engineers, AI researchers, and data center operations teams to co-design high-performance infrastructure.
- Provide technical leadership and mentorship to network engineers while establishing best practices and operational standards.
- Contribute to the long-term networking strategy and roadmap for Graphcore’s AI infrastructure.
- Research and evaluate next-generation high-speed networking technologies and vendor solutions.
Requirements
- BS or MS or equivalent experience in Computer Science, Electrical Engineering, Network Engineering, or a related technical discipline.
- 12+ years of progressive network engineering experience, with at least 3 years in hyperscale, high-density, or HPC data center environments.
- Expert-level knowledge of data center routing and switching protocols including BGP, OSPF, and EVPN-VXLAN architectures.
- Strong operational understanding of RDMA networking technologies such as RoCEv2 or InfiniBand.
- Hands-on experience with modern merchant silicon networking platforms and NOS platforms such as Arista EOS, Cisco NX-OS, or SONiC.
- Experience deploying high-speed network technologies including 400G/800G optics and large-scale fabric architectures.
- Proficiency in automation and scripting languages such as Python, Go, Bash, or similar tools.
- Strong collaboration and communication skills across cross-functional engineering teams.
Desirable Skills
- Experience operating large-scale AI or GPU clusters.
- Familiarity with network telemetry frameworks and streaming analytics.
- Experience implementing NetDevOps workflows and infrastructure automation pipelines.
- Experience influencing vendor roadmaps or evaluating next-generation networking technologies.
Benefits
- Medical, dental, and vision coverage.
- Flexible Spending Accounts (FSAs) and Health Savings Accounts (HSAs).
- Disability and life insurance.
- 401(k) retirement plan.
- Commuter benefits.
- Wellness services and an Employee Assistance Programme (EAP).