Jobs · Information Technology

Incident Management Specialist

The UVA VEC · United States · 6 days ago
RemoteRemoteInformation Technology$80–$120/hrOther

Role Overview

We are seeking experienced Incident Management, Reliability, and Site Reliability Engineering (SRE) professionals to support an advanced AI evaluation project. In this role, you will leverage your technical expertise to assess AI-generated technical documents, presentations, reports, and spreadsheets for accuracy, completeness, engineering quality, and adherence to industry best practices. Rather than building infrastructure or responding to incidents directly, you will evaluate AI-generated deliverables, provide structured technical feedback, and help improve next-generation AI systems used in reliability engineering and IT operations.

Key Responsibilities

  • Evaluate AI-generated documents, spreadsheets, presentations, and technical reports related to incident management, reliability engineering, and SRE.
  • Assess technical accuracy, operational best practices, clarity, completeness, and overall quality using structured evaluation rubrics.
  • Identify factual errors, technical inconsistencies, presentation issues, and gaps in engineering reasoning.
  • Provide detailed written feedback with clear, evidence-based justifications for evaluation decisions.
  • Review AI-generated recommendations involving incident response, service reliability, monitoring, observability, root cause analysis, and operational excellence.
  • Apply consistent engineering judgment to ensure reliable and reproducible evaluations.
  • Collaborate remotely with project teams while maintaining high standards of quality and attention to detail.

Preferred Qualifications

  • 5+ years of professional experience in Incident Management, Site Reliability Engineering (SRE), Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Operations.
  • Strong understanding of production systems, incident response processes, service monitoring, and operational reliability.
  • Experience performing root cause analysis (RCA), post-incident reviews, and reliability improvement initiatives.
  • Advanced proficiency with Microsoft Office and Google Workspace, particularly PowerPoint/Google Slides, Excel/Sheets, and technical documentation.
  • Excellent analytical thinking and technical writing skills.
  • Ability to evaluate complex engineering content with precision and consistency.
  • Master's degree or higher in Computer Science, Information Technology, Engineering, or a related discipline is preferred.

Preferred Skills

  • Site Reliability Engineering (SRE)
  • Incident Management & Response
  • Root Cause Analysis (RCA)
  • Reliability Engineering
  • DevOps & Infrastructure Operations
  • Monitoring & Observability
  • Service Level Objectives (SLOs) & SLAs
  • Technical Documentation Review
  • Microsoft Office & Google Workspace
  • AI Evaluation & Quality Assessment (preferred)

Why Join

  • Contribute to the development of next-generation AI systems for IT operations and reliability engineering.
  • Flexible, fully remote contractor opportunity.
  • Competitive hourly compensation.
  • Weekly payments.
  • Work alongside experienced engineering professionals and AI researchers.
  • No prior AI experience required—your domain expertise is what matters most.

Similar jobs