Jobs · Engineering

Reliability Engineer

Bright Vision Technologies · Santa Clara, CA · 3 wk ago
RemoteRemoteEngineering$75k–$95k/yrFull-time

About the role

We are seeking an experienced Reliability Engineer to ensure the availability, performance, and operational excellence of large-scale distributed systems in production. As an SRE you will live at the boundary between development and operations, applying strong software engineering principles to infrastructure and operations problems, and continually pushing the platform toward higher reliability with lower operational toil. The ideal candidate will combine deep systems knowledge with strong programming skills, a measurement-driven mindset, and the discipline to design, automate, and operate complex services so that reliability becomes a first-class engineering deliverable rather than a reactive concern.

Responsibilities

  • Support large-scale distributed systems in production
  • Apply software engineering principles to infrastructure and operations problems
  • Design, automate, and operate complex services to improve reliability
  • Lead incident response and conduct effective post-incident reviews

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related technical discipline
  • Five or more years of SRE, DevOps, or production engineering experience supporting large-scale distributed systems
  • Strong programming skills in at least one of Python, Go, or Java, with the ability to build robust automation and tooling
  • Deep, hands-on experience operating Linux at scale, including networking, performance tuning, and systems-level troubleshooting
  • Production experience operating Kubernetes and container-based workloads
  • Strong working knowledge of observability tooling such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, or commercial equivalents
  • Hands-on experience designing and operating CI/CD pipelines for both infrastructure and applications
  • Solid understanding of distributed system design, including consistency models, partitioning, and failure semantics
  • Excellent communication and documentation skills

Preferred Qualifications

  • Experience defining and operationalizing SLOs and error budgets in real production environments
  • Exposure to chaos engineering practices and tools such as Chaos Monkey, Gremlin, or Litmus
  • Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP)
  • Background in capacity planning, performance engineering, or large-scale load testing
  • Familiarity with service mesh technologies such as Istio, Linkerd, or Consul

Pay

$75,000–$95,000 Annually

U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.

Similar jobs

Reliability Engineer

Decision Technologies, Inc.Arlington, VA· 3 wk ago
Information Technologyapply on ats.rippling.com

Reliability Engineer

Pratt & WhitneyAsheville, NC· 1 mo ago
Information Technologyapply on careers.rtx.com

Reliability Engineer

SyngentaOmaha, NE· 2 mo ago
Supply Chain$250k/yrapply on jobs.smartrecruiters.com

Reliability Engineer

I-care Reliability Inc.Pennsylvania, United States· 1 mo ago
Engineeringapply on icare-usa.cvw.io

Reliability Engineer

James HardieUnited States· 1 mo ago
RemoteEngineering$125k–$145k/yrapply on careers.jameshardie.com