Cloud Problem Analyst
THUNDERYARD SOLUTIONS · United States · 4 wk ago
RemoteRemoteInformation Technology$110k–$120k/yrFull-time
Key Responsibilities
- Own the end-to-end ITIL Problem Management lifecycle for AWS and Azure platform services, from problem identification through closure and continuous improvement.
- Analyze incident trends, major incidents, monitoring data, service health events, and operational metrics to identify recurring or high-impact problems.
- Facilitate root cause analysis sessions with cloud engineering, infrastructure, application, network, database, security, and vendor teams.
- Document problem records, known errors, workarounds, corrective actions, and permanent resolutions.
- Drive completion of corrective and preventive actions, ensuring ownership, target dates, risk visibility, and measurable outcomes.
- Partner with Incident Management, Change Management, Release Management, SRE, DevOps, and platform operations teams to prevent recurrence of incidents.
- Develop and maintain problem management dashboards, KPIs, reports, and executive summaries for operational reviews.
- Evaluate service reliability patterns across AWS and Azure services, including compute, storage, networking, identity, monitoring, backup, and platform automation capabilities.
- Promote a blameless, fact-based problem management culture focused on learning, prevention, and service improvement.
- Contribute to continual service improvement initiatives that improve cloud service availability, resiliency, observability, and operational maturity.
Must-Have Qualifications
- 5+ years of IT operations, service management, infrastructure, cloud operations, or related technology experience.
- 3+ years of hands-on experience leading ITIL Problem Management in a complex enterprise environment.
- Strong understanding of ITIL practices, including Incident Management, Problem Management, Change Enablement, Knowledge Management, Service Level Management, and Continual Improvement.
- Demonstrated ability to lead root cause analysis using structured methods such as 5 Whys, fishbone diagrams, fault-tree analysis, Kepner-Tregoe, or similar approaches.
- Experience using ITSM tools such as ServiceNow, Jira Service Management, BMC Helix, or equivalent platforms to manage problem records and corrective actions.
- Able to interpret incident data, operational telemetry, service health alerts, and trend reports to identify systemic issues.
- Excellent written and verbal communication skills, including the ability to translate technical findings into business-oriented summaries.
- Strong stakeholder management skills with the ability to influence technical teams without direct authority.
- Experience coordinating cross-functional teams during post-incident reviews, problem review boards, and service improvement forums.
- Able to manage multiple high-priority problems simultaneously while maintaining quality, documentation, and follow-through.
- ITIL Foundation certification or equivalent practical ITSM experience.
Nice-to-Have Qualifications
- ITIL 4 Managing Professional, ITIL Practitioner, or advanced ITSM certification.
- Experience supporting cloud platform services in AWS and Azure, with working knowledge of cloud infrastructure, networking, identity, monitoring, automation, and availability concepts.
- AWS certification such as AWS Certified Cloud Practitioner, AWS Certified Solutions Architect, AWS Certified SysOps Administrator, or equivalent experience.
- Microsoft Azure certification such as Azure Fundamentals, Azure Administrator Associate, Azure Solutions Architect Expert, or equivalent experience.
- Experience working with observability and monitoring tools such as Azure Monitor, AWS CloudWatch, Datadog, Splunk, Dynatrace, New Relic, Grafana, or Prometheus.
- Familiarity with SRE practices, error budgets, service level indicators, service level objectives, and reliability engineering principles.
- Experience in highly regulated environments with audit, compliance, security, or risk management requirements.
- Knowledge of automation, CI/CD, infrastructure as code, configuration management, and cloud governance practices.
- Experience with major incident management, post-incident reviews, service restoration, and executive communications.
- Familiarity with Agile, Scrum, Kanban, DevOps, or product operating models.
- Experience developing problem management playbooks, maturity models, reporting frameworks, and process improvement roadmaps.
Compensation
The salary is budgeted at $110,000-$120,000 annually, plus benefits. ThunderYard offers benefits including medical, dental and vision insurance, 401k matching, PTO, certification reimbursement and more.