Incident Commander
PENN Interactive · United States · 1 mo ago
RemoteRemoteInformation Technology$90k–$135k/yrFull-time
About the role
The Incident Commander will work cross-functionally across engineering, and be the front line for incidents and working with Release Engineering to help prevent new events. This role is responsible for all incidents across various organizations within the company, including P1, P2, P3, and P4.
Responsibilities
- Classify and document all incidents and carry out support, assist, and drive all incidents, including investigation, hierarchical and technical escalation, diagnosis, and recovery, and root cause analysis.
- Drive improvements to service delivery and release processes based on disruption reports.
- Lead and enhance collaboration with other Incident Commanders, Customer Support, Application, and Engineering teams - cross-functional teams to lead real-time incident management.
- Develop and maintain key practical capabilities, collaborating with SRE teams and infrastructure teams to identify requirements and gaps resulting in downtime or blindspots.
- Recommend innovative solutions that enable the organization to deliver on its objectives and goals.
- Promote opportunities for continuous service improvements, manage and update Root Cause Analysis documentation, and lead SRE communications to stakeholders via email, Slack, and Teams in a timely manner.
- Manage and update Root Cause Analysis documentation, lead SRE communications to stakeholders via email, Slack, and Teams in a timely manner, and lead initiatives to promote JIRA Release Ticket management, quality, and alignment with Incident management communication supporting SLAs.
- Other duties as required.
Requirements
- Experience in a similar role or incident management role.
- Experience and understanding of containerization (Docker & Kubernetes preferred).
- Automation: Understanding of configuration management and infrastructure as code tools (Terraform, Ansible, Helm, etc.).
- Experience with a programming language.
- Comfortable within Linux environments and needs.
- Experience working with AWS, GCP, and on-premises environments.
- Ability to work independently and learn quickly with little supervision.
- Ability to handle multiple projects simultaneously.
- Willingness to drop everything and take on an ad-hoc task.
Qualifications
- Tech-savvy and passionate about learning new technologies and tools.
- Outgoing, and able to keep a conversation going naturally to extract needed information.
- A degree in computer science, engineering, or similar experience.
Skills
- Experience in incident management or similar roles.
- Containerization (Docker & Kubernetes).
- Automation (configuration management and infrastructure as code tools).
- Programming languages.
- Linux environments.
- Experience with AWS, GCP, and on-premises environments.
- Collaboration and leadership skills.
- Problem-solving and root cause analysis.
- Communication and stakeholder management.
Benefits
Competitive compensation package, fun, relaxed work environment, education and conference reimbursements, opportunities for career progression and mentoring others.