Senior Network Engineer
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers, ranging from AI researchers to enterprises and hyperscalers. Our mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence.
About the role
This position requires presence in our San Francisco, San Jose, or Bellevue office location 4 days per week; Lambda’s designated work-from-home day is currently Tuesday.
Responsibilities
- Help to build and scale Lambda's high-performance cloud network
- Deploy and configure networking hardware for new and existing clusters
- Ensure high availability of our network through monitoring, failover, and redundancy
- Contribute to automation of network configuration management and operation
- Work with internal and external customers to resolve network-related issues
- Deploy and maintain network monitoring and management tools
- Participate in day-2 operations and on-call rotation for the Network Engineering team
Requirements
- 10+ years of experience in IT and networking
- 6+ years of experience designing and operating production data-center networks
- Led the implementation of large production-scale networking projects
- Experience managing Next-Generation Firewalls (e.g., Fortigate)
- Experience with cloud providers networking (AWS, GCP, OCI)
- Expertise in CLOS/Spine and Leaf fabrics, EVPN/VXLAN, ECMP, BGP, and fast convergence techniques
- Comfortable on the Linux command line with an understanding of the Linux networking stack and internals
- Strong automation skills (Python, Ansible) and experience with network APIs and git or similar source control systems
- Production experience with multiple network gear vendors (Arista, Juniper, Cisco, Cumulus/SONiC, Opengear)
- Experience with Networking Monitoring stack (Datadog, Clickhouse, Grafana, Prometheus, gNMI, OTel)
Nice to Have
- Knowledge or experience maintaining Software Defined Networks (SDN)
- Experience automating network configuration within public clouds using tools like Terraform/Ansible/Salt
- Hands-on with HPC/AI networking: RoCEv2 and/or InfiniBand (Congestion Control, VLs, partitions), GPUDirect RDMA concepts
- Experience with DWDM technologies and SD-WAN
- Understanding of data center power/space/cooling trade-offs and their impact on topology choices
- Experience with virtualization technology, like ESXi, KVM, and VMs management
- Experience with LoadBalancers like F5, NetScaler
About Lambda
Founded in 2012, Lambda has 500+ employees and is growing fast. Our investors include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove. We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG. Our values are publicly available at https://lambda.ai/careers.
Pay
Compensation range: $250K - $370K annual salary.
Benefits
- Generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan that we all actually use