Jobs · Engineering · California

Senior Site Reliability Engineer

ServiceTitan · Glendale, CA · 3 days ago
Engineering$20k/yrFull-time

About the role

We are seeking a seasoned developer to lead the Infrastructure engineering team at ServiceTitan. The role involves enhancing, building, and scaling our services to meet the needs of thousands of businesses worldwide. We are particularly interested in individuals with a strong background in web application development, distributed systems, and a proven track record of delivering technical leadership.

Responsibilities

  • Design, develop, test, troubleshoot, debug, optimize, scale, perform capacity planning, deploy, maintain, and improve software applications.
  • Coding and automation of applications on cloud platforms.
  • Collaborate with Product Engineering teams to plan and deploy product releases.
  • Work with Engineering leadership to build scalable infrastructure and shared services.
  • Contribute to product development as needed to ensure high-quality service availability.
  • Resolve product/service defects, design changes, infrastructure changes, or operational changes.
  • Establish strong relationships with company leadership to ensure the use of well-suited technologies.

Requirements

  • Bachelor's or Master's degree in Computer Science (or equivalent diploma and/or certifications) with 8-10 years of related experience.
  • Familiarity with Continuous Integration and Delivery including experience with tools such as Jenkins or Team City.
  • Strong expertise & experience with Kubernetes, Functions/Serverless computing, Distributed messaging systems, Data Lakehouse architectures, and API gateways.
  • Hands-on experience with a Distributed Version Control System such as Git.
  • Advanced knowledge of at least one of the following programming languages: C#, Visual Basic, PowerShell, Java.
  • Experience scripting provisioning of servers, applications, and/or infrastructure in a production environment at scale.
  • Knowledge of software development best practices, SDLC, and experience deploying high availability systems and software.
  • Experience with troubleshooting distributed web applications in a production environment.
  • Experience with log / metric collection and analysis tools (e.g. Elasticsearch-Logstash-Kibana, DataDog, Grafana).

Qualifications

  • Administer and automate infrastructure for Azure, AWS, or other public clouds.
  • Subject matter expert in Cloud Infrastructure & Systems Reliability.
  • Comfortable owning a wide and diverse set of problem areas and willing to go out of your lane to affect change.
  • Developed one or more metrics, log aggregation, or performance analysis systems in your career.

Skills

  • Ability to operate within the trade-offs present when solving for immediate needs versus solving with bigger scale solutions.
  • Willingness to guide and make decisions with limited information.
  • Experience in a rapidly growing and changing environment.

Benefits

  • Flextime, recognition, and support for autonomous work.
  • Comprehensive onboarding program, leadership training, and other programs and events.
  • Great work is rewarded through Bonusly, peer-nominated awards, and more.
  • Holistic health and wellness benefits: medical, dental, vision, FSA, HSA, 401k match, telehealth options, and more.
  • Support for Titans at all stages of life: parental leave, fertility services, surrogacy, adoption reimbursement, and more.

Pay

Actual compensation within a range is determined by factors including relevant experience, skill set, qualifications, and performance. In addition to base salary, our total compensation package includes an annual bonus, equity, and a holistic suite of benefits.

Schedule

ServiceTitan offers flexible time off and ample learning and development opportunities to continue growing your career.

Similar jobs