Jobs · Engineering

Site Reliability Engineer - Production Support

Bayside Solutions · Cupertino, CA · 1 wk ago
RemoteRemoteEngineering$60–$70/hrContract

W2 contract role based in Cupertino, CA – fully remote.

About the role

Production support / application support / NOC / service reliability / incident management. Screen out security analysts whose Splunk work is log-based threat detection.

Responsibilities

  • Watch service health across production; catch, triage, and drive issues to resolution.
  • Run production support against SLAs and work the incident lifecycle directly with the incident management (IM) team.
  • Monitor and investigate using Splunk (and Datadog, where in play).
  • Support the application stack in cloud and datacenter environments.
  • Handle non-prod and prod deployments using the in-house CD tool; monitor deployments and roll back / debug when they fail.
  • Keep Git repos current; debug deployment scripts.

Requirements

  • Hands-on production support / application support in an SLA-bound environment. Must be able to describe what they personally did during a real incident.
  • Incident management experience: triage, severity assessment, escalation, bridge participation, driving to resolution, RCA follow-up. Direct interaction with an IM team.
  • Splunk is required. Candidates will be asked what commands / searches they run.
  • SOC / production monitoring.
  • Troubleshooting depth demonstrable at the command line: log tracing, connection failures, latency, service crashes. Conceptual answers fail here.
  • Linux / Unix fluency.
  • Containerized workloads: Docker + Kubernetes (managing, not just using).
  • Cloud: AWS primary; Azure / AKS acceptable and precedented.
  • Payments / banking production environment (card networks, Visa / Mastercard, wallets, issuer / acquirer processing).
  • Java application troubleshooting.

Preferred Qualifications

  • CI/CD tooling – specifically Jenkins; ability to debug the scripts behind a deployment, not just trigger one.
  • Python and / or shell scripting for automation and runbook work.
  • Load balancers, SSL/TLS and certificate management, DNS, and general network troubleshooting.
  • Config management (Ansible / Chef / Puppet), autoscaling design in Kubernetes.
  • Terraform / infrastructure-as-code exposure.
  • ServiceNow / Jira or equivalent ITSM incident tooling; ITIL familiarity.
  • Datadog experience is a plus.

Pay

$60 – $70 per hour.

Similar jobs