Senior Software Engineer, Release Infra
Brex · Seattle, WA · 3 wk ago
HybridEngineering$192k–$240k/yrFull-time
Responsibilities
- Design, build, and maintain the release infrastructure that powers Brex’s deployment pipelines and incident workflows
- Drive technical strategy and architecture for release and observability systems, making them more scalable, reliable, and secure
- Collaborate with product, engineering, and operations partners to ensure Brex’s releases are safe, predictable, and low-friction
- Identify and deliver improvements to the end-to-end release process (from code merge to production) to reduce risk and cycle time
- Build and evolve tooling for observability and incident response, enabling fast detection, triage, and resolution
- Proactively identify and mitigate risks in our release and infrastructure stack, including performance, reliability, and security concerns
- Define, instrument, and monitor key metrics for release engineering (e.g., deployment frequency, change failure rate, MTTR) and use them to guide improvements
- Partner with other infrastructure and product teams to debug complex production issues and drive long-term fixes
- Contribute to and champion best practices in release engineering, reliability, and operational excellence across the organization
- Mentor other engineers on the team, providing technical guidance and code reviews to elevate the overall quality of our infrastructure
- Stay up-to-date on emerging tools and practices in release engineering, observability, and SRE, and bring relevant ideas into Brex’s stack
Requirements
- 7+ years of professional experience designing, building, and operating backend or infrastructure systems in production
- Strong proficiency in backend programming languages (e.g., Go, Java, Kotlin, or Python) with a focus on reliability and performance
- Hands-on experience with CI/CD and release pipelines (e.g., GitHub Actions, CircleCI, Buildkite, Argo, Spinnaker, Jenkins) including build, test, and deployment automation
- Experience architecting and operating scalable, high-availability distributed systems on cloud platforms (e.g., AWS, GCP, Azure)
- Deep familiarity with containerization and orchestration (e.g., Docker, Kubernetes) and infrastructure-as-code (e.g., Terraform, CloudFormation)
- Experience designing and maintaining observability tooling (metrics, logs, tracing) and integrating it into incident response workflows
- Strong understanding of reliability and SRE practices, including SLIs/SLOs, error budgets, and incident management best practices
- Experience designing and optimizing data storage systems (SQL and/or NoSQL) for operational and observability use cases
- Proven track record of improving release processes (e.g., reducing deployment risk, increasing deployment frequency, automating rollbacks)
- Comfort working cross-functionally with product and other engineering teams to debug complex production issues and ship changes safely
- Strong communication and collaboration skills, including writing clear design docs and driving technical decisions across teams