Sr. Site Reliability Engineer
FreedomPay · Philadelphia, PA · 4 days ago
Full-time
About the role
The Senior Site Reliability Engineer will play a critical role in ensuring the highest possible availability and resiliency of FreedomPay’s global payment platform. This position is part of a global team that works closely with other engineering teams to solve complex problems and reduce manual tasks.
Primary Responsibilities
- Build and maintain a comprehensive understanding of the platform and custom application stack.
- Implement, maintain, and continuously improve observability strategies and metrics that ensure complete system health for numerous complex products throughout all stages of the development lifecycle, up to and including production.
- Continuously identify automation opportunities and follow through to successful implementation, applying AI-assisted tooling to accelerate development and reduce manual effort.
- Design, build, and maintain automated remediation and self-healing workflows that detect, triage, and resolve common failure modes with minimal human intervention.
- Leverage AI/ML-driven observability — anomaly detection, alert correlation, and intelligent noise reduction — to surface issues earlier and shorten time to detection.
- Use AI-assisted analysis to accelerate root-cause investigation, enrich incident context, and generate first-draft postmortems and runbooks for human review.
- Handle escalations and collaborate effectively with other team members to quickly determine the root cause of any type of service degradation.
- Implement, maintain, and continuously improve incident response procedures and other operational documentation, automating documentation generation and upkeep wherever practical.
- Aid with troubleshooting and remediation of failed scheduled jobs and data-related concerns.
- Champion responsible, secure adoption of AI tooling across the SRE function — sharing patterns, prompts, and automations that raise the productivity of the whole team.
Required Background and Experience
- BS degree in Computer Science or equivalent, or equivalent years of relevant experience.
- Minimum of 5 years of hands-on technical experience in highly available, high-throughput, web-based technology environments.
- Demonstrated history of self-directed learning.
- Next-level problem-solving abilities and a strong bias toward practical, proven solutions.
- Track record of identifying and eliminating manual toil through automation.
- Excellent communication and organizational skills, with a strong sense of ownership and service.
Required Technical Skills
- Expert-level proficiency in an enterprise APM platform and its AI/ML-driven (AIOps) capabilities; Dynatrace experience strongly preferred, though deep expertise in comparable tools such as Datadog or New Relic where readily transferable.
- Hands-on experience with AI-assisted development and automation tools — such as Anthropic (Claude), OpenAI (Codex), and Azure AI services (Foundry, Azure SRE Agent) — and a demonstrated ability to apply them to real operational and engineering work.
- Proficiency in scripting and automation — PowerShell and/or Python — to build tooling and remediation workflows.
- Strong SQL / T-SQL skills.
- Solid understanding of core networking concepts: DNS, HTTP/HTTPS, load balancing, and TCP/IP routing and switching.
- Working knowledge of modern technology infrastructure including container orchestration, IaaS/PaaS cloud services, Azure, and VMware.
- Working knowledge of application development processes.