Jobs · Information Technology · Washington

Staff Engineer, Network Observability

CoreWeave · Bellevue, WA · 3 days ago
Information Technology$207k–$275k/yrFull-time

CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025.

About the role

The Network Observability team is responsible for how CoreWeave observes, understands, and operates its network at scale. As a Staff Engineer for Network Observability, you will define and evolve the technical direction for network observability, partnering across Network Engineering, SRE, Platform, and adjacent infrastructure teams to build resilient telemetry systems, raise engineering standards, and turn observability into a strategic advantage for the business. Your mission: build and scale a network observability platform that provides CoreWeave fast, trustworthy insight into network behavior, enables proactive risk detection, improves how engineering teams make decisions during both normal operations and incidents, and enables closed-loop automation workflows.

Responsibilities

  • Set technical direction for network observability across multiple teams, ensuring the platform, data models, and telemetry strategy align with long-term engineering and business goals
  • Lead the design and evolution of scalable observability solutions using diverse collector technologies (e.g., gNMI, SNMP, Prometheus scraping, OTEL, logs, flows, etc.), persistence databases (e.g., Prometheus-like, Loki, Clickhouse), and visualization and alerting ones (e.g., Grafana, Alertmanager), with a strong focus on reliability, usability, and future scale
  • Drive cross-team initiatives to standardize observability patterns, improve signal quality, and create a consistent approach to logs, metrics, events, flows, and related diagnostics across the network stack
  • Partner closely with engineering leadership and technical stakeholders to prioritize investments, navigate ambiguity, and make high-leverage technical tradeoffs that improve resilience, scalability, and operator efficiency
  • Act as a go-to technical expert for critical observability challenges, especially when incidents, architectural complexity, or unclear ownership require strong judgment and coordination
  • Mentor junior and senior engineers through technical reviews, design guidance, and hands-on problem solving, raising the bar for engineering quality and multiplying the impact of the broader team
  • Participate in design discussions, RFCs, and architectural decisions across the broader infrastructure organization, helping teams converge on scalable, maintainable solutions
  • Join a rotating on-call schedule as a senior escalation point for observability-related issues, helping teams quickly isolate failures, improve incident response, and drive durable follow-through after outages

Requirements

  • Deep expertise in building flexible network observability solutions, with diverse implementation options for collectors, distribution, processing, persistence, alerting, analytics, and visualization
  • Experience as a Network Engineer, SRE, Software Engineer, or Systems Engineer in large-scale environments, with a strong track record of building and operating observability or infrastructure platforms that support multiple teams
  • Demonstrated ability to lead through ambiguity, shape technical direction, and make sound architectural and operational tradeoffs that balance immediate needs with long-term maintainability
  • Strong systems thinking and practical experience designing resilient, scalable solutions that improve visibility, incident response, and engineering efficiency
  • Proven ability to work across teams and functions, influence without formal authority, and build trust with both technical and non-technical stakeholders
  • Proficient with Python, Go, and Bash, plus familiarity with configuration management and templating tools such as Ansible and Jinja2
  • Comfortable containerizing and operating solutions in Kubernetes, including designing, building, and deploying container-based workloads efficiently
  • Strong knowledge of Linux systems and IP networking concepts, with hands-on experience in routing, switching, and network troubleshooting
  • Practical experience with networking platforms such as SONiC, HPE Junos, NVIDIA Cumulus Linux, Nokia SR OS, and SR Linux
  • A strong mentorship mindset and a history of helping other engineers grow through coaching, design feedback, documentation, and technical leadership

Preferred Qualifications

  • Bachelor's degree in Computer Science or a related field
  • Hands-on experience applying machine learning techniques or tools to proactively detect performance or security anomalies in network traffic
  • Experience with OpenTelemetry, Jaeger, Zipkin, or similar tooling for end-to-end tracing across distributed systems and infrastructure components
  • Experience shaping technical roadmaps, setting standards, or leading platform investments that materially improved reliability or scalability across multiple teams
  • Network certifications such as CCNA, CCNP, or similar

About CoreWeave

At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values:

  • Be Curious at Your Core
  • Act Like an Owner
  • Empower Employees
  • Deliver Best-in-Class Client Experiences
  • Achieve More Together

We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and enables the development of innovative solutions to complex problems. As we get set for takeoff, the organization's growth opportunities are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too.

Pay

The base salary range for this role is $207,000 to $275,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility).

Benefits

The range we’ve posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location. In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings; for roles in other locations, benefits vary and are shared during the hiring process.

  • Medical, dental, and vision insurance - 100% paid for by CoreWeave
  • Company-paid Life Insurance
  • Voluntary supplemental life insurance
  • Short and long-term disability insurance
  • Flexible Spending Account
  • Health Savings Account
  • Tuition Reimbursement
  • Ability to Participate in Employee Stock Purchase Program (ESPP)
  • Mental Wellness Benefits through Spring Health
  • Family-Forming support provided by Carrot
  • Paid Parental Leave
  • Flexible, full-service childcare support with Kinside
  • 401(k) with a generous employer match
  • Flexible PTO
  • Catered lunch each day in our office and data center locations
  • A casual work environment
  • A work culture focused on innovative disruption

Similar jobs