Jobs · Rhode Island

Only W2: SRE / Production Reliability Engineer, Woonsocket, RI

Tror - AI for everyone · Woonsocket, RI · 3 days ago
On-siteContract

What You'll Do

  • Own reliability and performance of critical production applications.
  • Act as an Incident Commander (IC) during P1/P2 production incidents.
  • Lead incident response, root-cause analysis, postmortems, and reliability improvements.
  • Define and manage SLIs, SLOs, and error budgets.
  • Build and improve monitoring, alerting, and observability solutions.
  • Tune and validate time-series anomaly detection models for production monitoring.
  • Develop automation and operational tools using Python, Java, and React.
  • Troubleshoot Kubernetes, cloud, batch processing, and data pipeline issues.
  • Work with engineering and operations teams to improve system reliability and reduce manual work.
  • Support large-scale deployments and manage production risks such as configuration drift and blast radius.

Must-Have Skills

  • 8+ years of experience in SRE, DevOps, Platform Engineering, or Production Engineering.
  • Hands-on experience as an Incident Commander for P1/P2 incidents.
  • Strong experience with time-series anomaly detection models in production observability (mandatory).
  • Strong production-level programming skills in: Python, Java, React.
  • Strong experience with SLI, SLO, and error budgets.
  • Hands-on observability experience with: Prometheus, Grafana, OpenTelemetry.
  • At least 2 log platforms such as Loki, Splunk, or Elasticsearch.
  • Strong GCP experience.
  • Strong Kubernetes operational experience.
  • Experience with Rancher K3s.
  • Experience troubleshooting Apache Airflow and Tidal workflows/batch jobs.
  • Experience with production-scale distributed systems and on-call support.

Preferred Skills

  • Production Readiness Reviews / service launch experience.
  • BigQuery and PostgreSQL.
  • Chaos Engineering / fault injection.
  • TIC/Technical Incident Commander certification.
  • Healthcare, pharmacy, retail, or other high-availability environments.
  • LLM/GenAI for incident management, alert summarization, or runbook recommendations.
  • Kafka.
  • Istio / Envoy.
  • Terraform / Ansible.

Similar jobs

Reliability Engineer, Sr.

Bastion Technologies, Inc.Houston, TX· 2 mo ago
Engineeringapply on bastiontechnologies.applicantpro.com