Platform Operations Engineer
FreedomPay · Philadelphia, PA · 3 wk ago
HybridFull-time
Primary Responsibilities
- Identify monitoring and performance optimization opportunities, implement improvements, and continuously optimize system monitoring.
- Conduct root-cause analysis for platform incidents, implement and maintain monitoring tools, and develop proactive strategies and solutions to prevent recurrence.
- Serve as an escalation point for complex technical issues related to platform and services performance and stability.
- Manage problems and coordinate escalation and resolution efforts, developing solutions that minimize business impact.
- Document procedure improvements and maintain detailed SOP and ARP documentation to support operational consistency and knowledge retention.
- Continuously build technical knowledge and act as a subject-matter expert across multiple platform areas, while maintaining strong relationships with stakeholders across FreedomPay.
Required Background And Experience
- BS degree in Computer Science or equivalent, or equivalent years of relevant experience.
- 4+ years of hands-on technical experience in highly available, high-throughput, web-based or transaction-processing environments.
- Strong problem-solving skills, with a track record of researching and applying the latest technological strategies to improve platform operations.
- Strong incident management, with a track record of driving issues to resolution and minimizing business impact.
- Excellent communication and organizational skills, with a strong sense of ownership and service.
Required Technical Skills
- Proficiency with an enterprise APM or observability platform (e.g., Dynatrace, Datadog, New Relic, or comparable) for monitoring and performance optimization.
- Experience with AIOps and ML-driven observability capabilities, including anomaly detection, alert correlation, and intelligent noise reduction, to surface issues earlier and shorten time to detection.
- Experience with log management and analysis tools such as Splunk or similar platforms.
- Experience with AI-assisted development and operations tools such as Claude (Anthropic) or similar platforms.
- Proficiency in scripting and automation, including PowerShell and/or Python, to support monitoring, tooling, and remediation work.
- Solid understanding of core networking concepts: firewalls, load balancing, and TCP/IP routing and switching.
- Working knowledge of modern technology infrastructure, including Kubernetes, Terraform, IaaS/PaaS cloud services, Azure, and VMware.
- Experience with Azure Logic Apps, Azure Functions, or Azure AI Foundry for automation and AI-driven workflows.
- Strong SQL / T-SQL skills.
Preferred Qualifications
- Experience with payment processing systems or financial services.
- Knowledge of PCI policies and best practices.
- Experience mentoring and sharing knowledge with team members.