Jobs · Engineering

Software Engineer, Site Reliability

Upstart · United States · 6 days ago
RemoteRemoteEngineering$142k–$197k/yrFull-time

At Upstart, we’re united by a mission that matters: to radically reduce the cost and complexity of borrowing for all Americans. Every day, we bring creativity, experimentation, and advanced AI to reshape access to credit, helping millions move forward financially with clarity and confidence. As the leading AI lending marketplace, we partner with banks and credit unions to expand access to affordable credit through technology that’s both radically intelligent and deeply human. Our platform runs over one million predictions per borrower using more than 3,000 signals, powering smarter, fairer decisions for millions of customers.

We’re proudly digital-first, giving most Upstarters the flexibility to do their best work from wherever they thrive, alongside teammates across 80+ cities in the US and Canada. Digital-first doesn’t mean distant. We’re intentional about in-person connection through team onsites, planning sessions, and moments that spark creativity and trust. Whether you choose to work primarily from home or collaborate in-person from one of our offices in Columbus, Austin, the Bay Area, or New York City (opening Summer 2026), you’ll have the support to work in the way that works best for you.

About the Team

Upstart’s Site Reliability Engineering team enables engineers to operate reliable, resilient, and observable production systems at scale. We build the platforms, tooling, automation, and operational practices that help teams understand system health, respond effectively when things go wrong, and continuously improve the reliability of the services they own. Our goal is to make reliability an integrated part of how software operates at Upstart.

We provide shared observability and reliability capabilities, improve incident response and operational readiness, automate recurring operational work, and use system and customer data to identify where reliability investments will have the greatest impact. SRE partners closely with product engineering, Cloud Platform, Delivery, Developer Platform, Security, and other infrastructure teams. SRE provides shared reliability capabilities and operational practices, while engineering teams remain accountable for the reliability and operation of the services they build.

Responsibilities

  • Build and improve the tooling, services, and automation that help engineers understand and improve the reliability of production systems
  • Develop shared observability capabilities that make metrics, logs, traces, service health, and customer impact easier to understand and act on
  • Improve incident response and operational readiness through better tooling, automation, standards, and actionable production signals
  • Build resiliency capabilities that help teams identify failure modes, reduce operational risk, and recover effectively from infrastructure or application failures
  • Identify recurring operational toil and reliability problems and replace manual processes with durable software and automation
  • Use AI as an integrated part of software development and operational problem solving, while identifying opportunities for AI enabled capabilities that improve incident investigation, observability, reliability, and engineering efficiency

Requirements

  • 3+ years of professional experience in software engineering, site reliability engineering, or a related engineering discipline
  • Strong software development skills in one or more general purpose programming languages such as Python, Go, JavaScript, or TypeScript
  • Experience designing, building, testing, and operating production software, internal tooling, or infrastructure
  • Experience with cloud infrastructure, distributed systems, observability, monitoring, or production operations
  • Experience participating in on call or incident response for production systems
  • Demonstrated ability to independently deliver well scoped engineering projects, navigate technical ambiguity, and collaborate effectively across engineering teams
  • Demonstrated experience using AI assisted development tools across multiple stages of the software engineering lifecycle, with an interest in continually evolving how you use these tools as their capabilities advance

Preferred Qualifications

  • Experience with Kubernetes, AWS, infrastructure as code, and cloud native production environments
  • Experience building internal reliability, observability, incident management, or operational automation tools
  • Experience with observability platforms such as Datadog, Sumo Logic, CloudWatch, or similar technologies
  • Experience with reliability practices such as service level objectives, capacity planning, resiliency testing, disaster recovery, or operational readiness
  • Experience operating distributed applications with complex dependencies and high availability requirements
  • Experience building AI enabled operational workflows, tools, or automation that extend beyond individual code generation

Location

This role is available remotely in the United States.

Travel Requirements

As a digital-first company, most work can be accomplished remotely. Employees are encouraged to participate in regular onsites, typically once or twice per quarter for 2-4 consecutive days at a time.

Pay

Anticipated base salary range for this position is $142,000 - $196,600 USD, depending on geographic location. Individual pay is determined by job-related skills, experience, and relevant education or training. In addition, Upstart provides target bonuses, equity compensation, and generous benefits packages.

Benefits

  • Competitive compensation, including base pay, bonus opportunities, and annual equity grants that vest quarterly
  • 401(k) or Group Retirement Savings Plan with a company match of $2 for every $1 contributed, up to $15,000 annually
  • Employee Stock Purchase Plan (ESPP) with discounted stock purchase options (US only)
  • Comprehensive health coverage, including medical, dental, vision, and wellness resources (US and supplemental for Canada)
  • Health Savings Account contributions from Upstart for eligible plans (US only)
  • Income protection benefits, including life insurance and disability coverage
  • Paid time off, sick leave, and company holidays
  • Paid family and parental leave (duration varies by country)
  • Family-centered benefits to support fertility, parenthood, and caregiving needs
  • Employee Assistance Program (EAP) offering mental health support and life-centered resources
  • Financial wellness resources, including access to financial planning tools and a financial concierge service (US only)
  • Annual wellness allowance to support physical and emotional well-being and personal development
  • Annual productivity allowance to invest in relevant tools and resources
  • Connection and community through team events, all-company updates, and employee resource groups (ERGs)
  • Onsite perks, including catered lunches and fully stocked micro-kitchens in offices (Bay Area, Austin, Columbus, and New York City)

Similar jobs