Jobs · Utah

Platform Site Reliability Engineer

HybridFull-time

About the role

Software Guidance & Assistance, Inc. (SGA) is seeking a Platform Site Reliability Engineer for a hybrid contract assignment with a premier SaaS client in Lehi, UT. This role focuses on ensuring the reliability, availability, and operational health of a portfolio of SaaS-based solutions—including both in-house (customer-zero) platforms and vendor-managed services—by applying Site Reliability Engineering principles in environments where full-stack control is not always possible. The position requires strong observability practices, operational judgment, and effective collaboration across teams and vendors, with a key emphasis on on-call operations, incident response, and driving practical improvements that reduce risk and customer impact. Success depends on communication, creative instrumentation and monitoring, and influencing reliability outcomes across organizational and vendor boundaries using modern, AI-assisted tooling.

Responsibilities

  • Ensure the reliability, availability, and operational health of a portfolio of SaaS-based solutions, including vendor-managed services and in-house (customer-zero) platforms.
  • Participate in on-call rotations and incident response, leading investigation, mitigation, coordination, and post-incident follow-up.
  • Establish and maintain effective observability for systems that are not fully owned, identifying practical ways to obtain actionable metrics, logs, and signals from vendor and partner solutions.
  • Use operational data and incident learnings to identify reliability risks and drive targeted improvements that reduce customer impact.
  • Apply appropriate change controls at owned or influenced layers of the stack, balancing reliability, velocity, and business needs.
  • Partner with internal teams and external vendors to communicate expectations, coordinate response and remediation, and influence reliability outcomes.
  • Produce clear incident communications and post-incident analyses that inform stakeholders and drive lasting improvements.
  • Leverage automation and AI-assisted tooling to improve detection, triage, and operational efficiency.

Required Skills

  • Strong foundation in Site Reliability Engineering practices, including observability, incident response, and reliability measurement.
  • Hands-on experience operating SaaS or third-party systems where full-stack ownership is limited.
  • Deep understanding of monitoring, logging, and alerting, with the ability to design signals that are actionable rather than noisy.
  • Proven incident response experience, including on-call participation and cross-team coordination during high-impact events.
  • Ability to think creatively and pragmatically when instrumenting and improving systems with constrained control.
  • Excellent written and verbal communication skills, especially in high-pressure incident and vendor-coordination scenarios.
  • Experience working across organizational and vendor boundaries to resolve complex operational issues.
  • Sound engineering judgment when assessing risk, prioritizing work, and making reliability tradeoffs in production environments.
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent practical experience.
  • 3–5 years of experience in Site Reliability Engineering, production operations, or a closely related role.
  • Experience supporting production systems with on-call responsibilities and incident response expectations.
  • Strong experience working with observability data (metrics, logs, alerts) to diagnose issues and drive improvements.
  • Comfort using automation and AI-assisted tools as part of everyday operational workflows.

Preferred Skills

  • Experience supporting enterprise-scale SaaS platforms or shared services.
  • Prior experience working directly with vendors to resolve reliability or operational issues.
  • Familiarity with cloud-based and distributed system architectures.

Similar jobs