Production Support
Zennsoft,Inc · New York, NY · 5 days ago
On-siteManagement$45/hrFull-time
Location: Buffalo, NY – Contract (C2C) at $45/hr
What You’ll Do
- Service support & incident management
- Monitor and support production services, responding to incidents and service requests within agreed SLAs.
- Perform initial triage (logs, monitoring, data checks), identify likely causes, and implement workarounds/fixes within your control.
- Coordinate recovery activities with partner teams (engineering, infrastructure, network, vendors).
- Escalate high-severity incidents to the Production Support Lead/Major Incident Manager.
- Provide timely, accurate updates on impact, recovery progress, and next steps to stakeholders.
- Problem management & continuous improvement
- Contribute to root cause analysis and post-incident reviews, capturing actions to prevent recurrence.
- Create and maintain runbooks, knowledge articles, and operational checklists to reduce time to restore service.
- Identify opportunities for automation and improved monitoring/alerting to reduce noise and improve resilience.
- Change & release support
- Participate in change and release activities (e.g., patching/evergreening, DR tests, role swaps), ensuring operational readiness.
- Review changes for supportability and operational risk; raise concerns to the Support Lead/Manager where appropriate.
- Toil reduction
- Review and contribute to avoiding repeated alerts from the production environment via defect management in partnership with the development team or optimizing the alert system.
- Automate daily manual routine work.
- Develop tools to simplify daily operational work.
Requirements
- Experience in production support / IT service management in a regulated environment, ideally within banking technology.
- Strong understanding of IT infrastructure concepts (on-prem and cloud), including networking fundamentals and service dependencies.
- Hands-on capability with Unix/Linux, SQL, Kubernetes, log analysis, and troubleshooting in distributed systems.
- Experience using monitoring/observability tools (e.g., Splunk, AppDynamics, Prometheus, Patrol, or equivalents).
- Working knowledge of Incident, Problem, Change, and Release Management (ITIL-aligned).
- Strong communication skills in English, comfortable working with global teams across time zones.
- Proven ability to collaborate across boundaries, prioritize effectively, and stay calm under pressure.