Client Services Specialist - Senior (U.S. Citizens Only) / Platform & Cloud
PCI Energy Solutions · United States · 4 days ago
RemoteRemoteOTHRFull-time
Responsibilities
- Administer the PCI application suite within hosted (cloud-based) environments, monitoring application and infrastructure health using tools such as Datadog, Grafana, and ELK (Kibana), and responding to alerts related to performance, capacity, and availability
- Triage and resolve incidents across production and lower environments, including monitoring support channels/calls and prioritizing incoming requests; leverage knowledge of the PCI application suite and hosted environments to diagnose issues, guide resolution, and determine appropriate escalation
- Act as a primary point of contact for operational issues and requests, coordinating across engineering, product operations, and client-facing teams to drive resolution, support system changes, and validate system performance following deployments or maintenance
- Execute routine operational tasks to maintain system stability, including deployments, patching, failover activities, and environment refresh processes (e.g., DR/HA, gold copies)
- Support cloud-based environments in Amazon Web Services (AWS), including compute, storage, and networking components, as well as application servers and database systems (e.g., Oracle)
- Follow and contribute to operational runbooks and standard procedures to ensure consistency and reliability in system operations
- Identify opportunities to improve efficiency by reducing manual processes and contributing to automation, monitoring enhancements, and alert tuning
- Manage and update work through ticketing systems (e.g., JIRA, ServiceNow), ensuring clear documentation and adherence to SLAs
Qualifications
- 5+ years of experience in technical operations, system administration, DevOps, or related roles
- Foundational knowledge of cloud platforms, preferably Amazon Web Services (AWS)(AWS Cloud Practitioner or Solutions Architect Associate certification is a plus)
- Understanding of core infrastructure concepts, including: Compute, networking, and storage
- Linux and/or Windows system administration
- Basic scripting (e.g., Python, Bash, or PowerShell)
- Exposure to or interest in site reliability and operational best practices, including monitoring, alerting, and incident response
- Ability to troubleshoot technical issues across systems, logs, and environments using a structured, analytical approach
- Familiarity with observability and monitoring tools such as: Datadog, Grafana, ELK stack (Kibana)
- Strong problem-solving mindset with a focus on automation, efficiency, and reducing manual processes
- Ability to manage multiple tasks, incidents, and priorities in a fast-paced operational environment
- Strong communication skills, including the ability to translate technical issues for non-technical stakeholders
- Experience with ticketing systems (e.g., JIRA, ServiceNow) and structured operational workflows is a plus
- Exposure to databases (e.g., Oracle) or enterprise application environments is a plus
- Prior experience in a client-facing or support role is a plus, particularly when combined with technical troubleshooting responsibilities