Jobs · Information Technology · Texas

Engineer, Site Reliability

T-Mobile · Frisco, TX · 3 days ago
Information TechnologyFull-time

At T-Mobile, we invest in YOU! Our Total Rewards Package ensures employees receive competitive compensation and wealth-building opportunities, including an annual stock grant, employee stock purchase plan, 401(k), and access to free, year-round money coaches.

About the role

This role is essential for maintaining and improving the reliability and resilience of digital infrastructure systems. It primarily involves automating processes, monitoring system health, and managing incident responses to reduce operational disruptions. Success is measured by system uptime, reduction in manual interventions, and rapid recovery from incidents. The work directly supports organizational service quality and operational performance by ensuring robust and reliable digital operations.

Responsibilities

  • Automate processes to improve system reliability and reduce manual operational tasks
  • Monitor systems proactively to minimize operational incidents and maintain service continuity
  • Streamline software development and deployment processes to enhance operational efficiency
  • Develop scripts and tools to decrease manual efforts in routine operational activities
  • Manage incident response to ensure rapid recovery and minimize service disruption
  • Adapt to new technologies to sustain and improve system robustness and performance
  • Perform other duties/projects as assigned by business management

Requirements

  • Bachelor's Degree plus 3 years of related work experience OR advanced degree with 1 year of related work experience OR combination of education and experience deemed equivalent (Required)
  • Acceptable areas of study include Computer Science or Engineering
  • Master's/Advanced Degree in Computer Science or Data Science (Preferred)
  • 2-4 years developing and maintaining CI/CD pipelines for software deployment (Preferred)
  • 2-4 years implementing and managing cloud-native platforms and solutions (Preferred)
  • 2-4 years guiding and mentoring teams in reliability engineering practices (Preferred)
  • At least 18 years of age
  • Legally authorized to work in the United States

Skills

  • Application Monitoring (Required)
  • Automation (Required)
  • CI/CD (Required)
  • Capacity Planning (Required)
  • Cloud Computing (Required)
  • Incident Management (Required)
  • Performance Tuning (Required)
  • Scripting (Required)
  • System Reliability (Required)

Preferred Certifications

  • Certified Kubernetes Administrator (CKA) – validates ability to use Kubernetes for automating deployment, scaling, and operations of application containers
  • AWS Certified DevOps Engineer – demonstrates expertise in provisioning, operating, and managing distributed application systems on AWS
  • Site Reliability Engineering (SRE) Foundation Certification – provides foundational understanding of SRE philosophy, practices, and tools

Pay

Base pay range: $84,900 - $153,200. The successful candidate’s actual pay will be based on work location, qualifications, and experience. Most Corporate employees are eligible for an annual bonus target of 15% based on company and/or individual performance.

Benefits

  • Medical, dental, and vision insurance
  • Flexible spending account
  • 401(k) with company match
  • Annual stock grant and employee stock purchase plan
  • Free, year-round money coaches
  • Paid time off and up to 12 paid holidays (about 4 weeks for new full-time employees, 2.5 weeks for new part-time employees annually)
  • Paid parental and family leave
  • Family building benefits
  • Back-up care and enhanced family support
  • Childcare subsidy
  • Tuition assistance and college coaching
  • Short- and long-term disability
  • Voluntary AD&D, accident, life, disability, and long-term care insurance
  • Mobile service and home internet discounts
  • Pet insurance
  • Commuter and transit programs

Similar jobs