Staff Software Engineer, Observability
Robinhood · Menlo Park, CA · 4 wk ago
On-siteEngineering$124/hrFull-time
About the team + role
We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards.
What you'll do
- Define and execute the full-stack observability roadmap, establishing a clear technical vision for metrics, logs, traces, and alerting infrastructure across Robinhood's engineering organization.
- Own and evolve the observability control plane, including telemetry ingestion pipelines, cost management strategies, and the tooling that ensures observability components remain highly available.
- Lead and collaborate with a team of six engineers to build scalable, self-service observability solutions that enable product and infrastructure teams to move faster with greater confidence.
- Establish and maintain Service Level Objectives (SLOs) for observability systems that meet or exceed Robinhood's 99.9% uptime target, ensuring the observability platform is as reliable as the services it monitors.
- Partner with the Robinhood Command Center and engineering teams across the organization to align on dependency mapping, incident response workflows, and observability standards.
What you bring
- 8+ years of software engineering experience, with a proven track record of owning and delivering large-scale observability or infrastructure platform initiatives.
- Deep expertise in Kubernetes and public cloud environments (AWS preferred), with the ability to architect and operate distributed systems at scale regardless of specific vendor tooling.
- Strong coding proficiency in one or more languages (Go, Python, or similar) with experience integrating observability agents, libraries, and instrumentation directly into production codebases.
- Demonstrated experience owning an observability control plane or telemetry pipeline — including ingestion cost management, cardinality control, and signal routing — in a high-traffic production environment.
- Experience with the Vector data pipeline (or equivalent high-throughput log/metrics routing tools) and familiarity with tools such as Prometheus, Grafana, Honeycomb, Humio, or Sentry is a plus.
What we offer
- Challenging, high-impact work to grow your career
- Performance driven compensation with multipliers for outsized impact, bonus programs, equity ownership, and 401(k) matching
- Top Tier benefits to fuel your work, including 100% paid health insurance for employees with 90% coverage for dependents
- Access to the best AI tools on the market and continuous AI skill-building for every employee, technical or not
- Lifestyle wallet - a highly flexible benefits spending account for wellness, learning, and more
- Employer-paid life & disability insurance, fertility benefits, and mental health benefits
- Time off to recharge including company holidays, paid time off, sick time, parental leave, and more!