Jobs · Information Technology

Cloud Problem Analyst

THUNDERYARD SOLUTIONS · United States · 4 wk ago
RemoteRemoteInformation Technology$110k–$120k/yrFull-time

Key Responsibilities

  • Own the end-to-end ITIL Problem Management lifecycle for AWS and Azure platform services, from problem identification through closure and continuous improvement.
  • Analyze incident trends, major incidents, monitoring data, service health events, and operational metrics to identify recurring or high-impact problems.
  • Facilitate root cause analysis sessions with cloud engineering, infrastructure, application, network, database, security, and vendor teams.
  • Document problem records, known errors, workarounds, corrective actions, and permanent resolutions.
  • Drive completion of corrective and preventive actions, ensuring ownership, target dates, risk visibility, and measurable outcomes.
  • Partner with Incident Management, Change Management, Release Management, SRE, DevOps, and platform operations teams to prevent recurrence of incidents.
  • Develop and maintain problem management dashboards, KPIs, reports, and executive summaries for operational reviews.
  • Evaluate service reliability patterns across AWS and Azure services, including compute, storage, networking, identity, monitoring, backup, and platform automation capabilities.
  • Promote a blameless, fact-based problem management culture focused on learning, prevention, and service improvement.
  • Contribute to continual service improvement initiatives that improve cloud service availability, resiliency, observability, and operational maturity.

Must-Have Qualifications

  • 5+ years of IT operations, service management, infrastructure, cloud operations, or related technology experience.
  • 3+ years of hands-on experience leading ITIL Problem Management in a complex enterprise environment.
  • Strong understanding of ITIL practices, including Incident Management, Problem Management, Change Enablement, Knowledge Management, Service Level Management, and Continual Improvement.
  • Demonstrated ability to lead root cause analysis using structured methods such as 5 Whys, fishbone diagrams, fault-tree analysis, Kepner-Tregoe, or similar approaches.
  • Experience using ITSM tools such as ServiceNow, Jira Service Management, BMC Helix, or equivalent platforms to manage problem records and corrective actions.
  • Able to interpret incident data, operational telemetry, service health alerts, and trend reports to identify systemic issues.
  • Excellent written and verbal communication skills, including the ability to translate technical findings into business-oriented summaries.
  • Strong stakeholder management skills with the ability to influence technical teams without direct authority.
  • Experience coordinating cross-functional teams during post-incident reviews, problem review boards, and service improvement forums.
  • Able to manage multiple high-priority problems simultaneously while maintaining quality, documentation, and follow-through.
  • ITIL Foundation certification or equivalent practical ITSM experience.

Nice-to-Have Qualifications

  • ITIL 4 Managing Professional, ITIL Practitioner, or advanced ITSM certification.
  • Experience supporting cloud platform services in AWS and Azure, with working knowledge of cloud infrastructure, networking, identity, monitoring, automation, and availability concepts.
  • AWS certification such as AWS Certified Cloud Practitioner, AWS Certified Solutions Architect, AWS Certified SysOps Administrator, or equivalent experience.
  • Microsoft Azure certification such as Azure Fundamentals, Azure Administrator Associate, Azure Solutions Architect Expert, or equivalent experience.
  • Experience working with observability and monitoring tools such as Azure Monitor, AWS CloudWatch, Datadog, Splunk, Dynatrace, New Relic, Grafana, or Prometheus.
  • Familiarity with SRE practices, error budgets, service level indicators, service level objectives, and reliability engineering principles.
  • Experience in highly regulated environments with audit, compliance, security, or risk management requirements.
  • Knowledge of automation, CI/CD, infrastructure as code, configuration management, and cloud governance practices.
  • Experience with major incident management, post-incident reviews, service restoration, and executive communications.
  • Familiarity with Agile, Scrum, Kanban, DevOps, or product operating models.
  • Experience developing problem management playbooks, maturity models, reporting frameworks, and process improvement roadmaps.

Compensation

The salary is budgeted at $110,000-$120,000 annually, plus benefits. ThunderYard offers benefits including medical, dental and vision insurance, 401k matching, PTO, certification reimbursement and more.

Similar jobs

Cloud Analyst

MaximusUnited States· 1 mo ago
RemoteInformation Technology$50k/yrapply on maximus.avature.net

Cloud Analyst

MEDITECHMassachusetts, United States· 2 mo ago
Information Technology$54k–$72k/yrapply on recruiting.paylocity.com

Cloud Analyst

SiloSmashers, Inc.Arlington, VA· 13 mo ago
Information Technologyapply on applicantpro.com

Cloud Systems Analyst

Texas Health and Human ServicesAustin, TX· 2 wk ago
Information Technology$6k–$9k/moapply on careers.hhs.texas.gov

Cloud Systems Analyst

Pedigo Staffing ServicesAustin, TX· 2 mo ago
Information Technologyapply on pedigostaffing.catsone.com