Jobs · Pennsylvania

Sr. Site Reliability Engineer

FreedomPay · Philadelphia, PA · 4 days ago
Full-time

About the role

The Senior Site Reliability Engineer will play a critical role in ensuring the highest possible availability and resiliency of FreedomPay’s global payment platform. This position is part of a global team that works closely with other engineering teams to solve complex problems and reduce manual tasks.

Primary Responsibilities

  • Build and maintain a comprehensive understanding of the platform and custom application stack.
  • Implement, maintain, and continuously improve observability strategies and metrics that ensure complete system health for numerous complex products throughout all stages of the development lifecycle, up to and including production.
  • Continuously identify automation opportunities and follow through to successful implementation, applying AI-assisted tooling to accelerate development and reduce manual effort.
  • Design, build, and maintain automated remediation and self-healing workflows that detect, triage, and resolve common failure modes with minimal human intervention.
  • Leverage AI/ML-driven observability — anomaly detection, alert correlation, and intelligent noise reduction — to surface issues earlier and shorten time to detection.
  • Use AI-assisted analysis to accelerate root-cause investigation, enrich incident context, and generate first-draft postmortems and runbooks for human review.
  • Handle escalations and collaborate effectively with other team members to quickly determine the root cause of any type of service degradation.
  • Implement, maintain, and continuously improve incident response procedures and other operational documentation, automating documentation generation and upkeep wherever practical.
  • Aid with troubleshooting and remediation of failed scheduled jobs and data-related concerns.
  • Champion responsible, secure adoption of AI tooling across the SRE function — sharing patterns, prompts, and automations that raise the productivity of the whole team.

Required Background and Experience

  • BS degree in Computer Science or equivalent, or equivalent years of relevant experience.
  • Minimum of 5 years of hands-on technical experience in highly available, high-throughput, web-based technology environments.
  • Demonstrated history of self-directed learning.
  • Next-level problem-solving abilities and a strong bias toward practical, proven solutions.
  • Track record of identifying and eliminating manual toil through automation.
  • Excellent communication and organizational skills, with a strong sense of ownership and service.

Required Technical Skills

  • Expert-level proficiency in an enterprise APM platform and its AI/ML-driven (AIOps) capabilities; Dynatrace experience strongly preferred, though deep expertise in comparable tools such as Datadog or New Relic where readily transferable.
  • Hands-on experience with AI-assisted development and automation tools — such as Anthropic (Claude), OpenAI (Codex), and Azure AI services (Foundry, Azure SRE Agent) — and a demonstrated ability to apply them to real operational and engineering work.
  • Proficiency in scripting and automation — PowerShell and/or Python — to build tooling and remediation workflows.
  • Strong SQL / T-SQL skills.
  • Solid understanding of core networking concepts: DNS, HTTP/HTTPS, load balancing, and TCP/IP routing and switching.
  • Working knowledge of modern technology infrastructure including container orchestration, IaaS/PaaS cloud services, Azure, and VMware.
  • Working knowledge of application development processes.

Similar jobs