Site Reliability Engineer - Production Support
Bayside Solutions · Cupertino, CA · 1 wk ago
RemoteRemoteEngineering$60–$70/hrContract
W2 contract role based in Cupertino, CA – fully remote.
About the role
Production support / application support / NOC / service reliability / incident management. Screen out security analysts whose Splunk work is log-based threat detection.
Responsibilities
- Watch service health across production; catch, triage, and drive issues to resolution.
- Run production support against SLAs and work the incident lifecycle directly with the incident management (IM) team.
- Monitor and investigate using Splunk (and Datadog, where in play).
- Support the application stack in cloud and datacenter environments.
- Handle non-prod and prod deployments using the in-house CD tool; monitor deployments and roll back / debug when they fail.
- Keep Git repos current; debug deployment scripts.
Requirements
- Hands-on production support / application support in an SLA-bound environment. Must be able to describe what they personally did during a real incident.
- Incident management experience: triage, severity assessment, escalation, bridge participation, driving to resolution, RCA follow-up. Direct interaction with an IM team.
- Splunk is required. Candidates will be asked what commands / searches they run.
- SOC / production monitoring.
- Troubleshooting depth demonstrable at the command line: log tracing, connection failures, latency, service crashes. Conceptual answers fail here.
- Linux / Unix fluency.
- Containerized workloads: Docker + Kubernetes (managing, not just using).
- Cloud: AWS primary; Azure / AKS acceptable and precedented.
- Payments / banking production environment (card networks, Visa / Mastercard, wallets, issuer / acquirer processing).
- Java application troubleshooting.
Preferred Qualifications
- CI/CD tooling – specifically Jenkins; ability to debug the scripts behind a deployment, not just trigger one.
- Python and / or shell scripting for automation and runbook work.
- Load balancers, SSL/TLS and certificate management, DNS, and general network troubleshooting.
- Config management (Ansible / Chef / Puppet), autoscaling design in Kubernetes.
- Terraform / infrastructure-as-code exposure.
- ServiceNow / Jira or equivalent ITSM incident tooling; ITIL familiarity.
- Datadog experience is a plus.
Pay
$60 – $70 per hour.