Jobs · New Jersey

Site Reliability Engineer (SRE), Data Products

CentralReach · Holmdel, NJ · Today
$135k–$160k/yrFull-time

Key Accountabilities

  • Support deployment, configuration, and daily operations across several AWS accounts and the Linux servers and workloads running in those environments.
  • Administer, monitor, and troubleshoot Amazon S3 buckets, AWS Glue jobs, Amazon Kinesis streams, AWS PrivateLink connectivity, security groups, and related AWS networking and access dependencies.
  • Investigate and resolve infrastructure incidents, deployment failures, service interruptions, and connectivity issues, with a focus on restoring service quickly and preventing recurrence.
  • Maintain and improve repeatable deployment and operational processes for AWS infrastructure and workloads through automation, version-controlled configuration, and clear runbooks.
  • Support secure connectivity to Snowflake accounts, including troubleshooting network paths, private connectivity, endpoint configuration, and environment-specific access issues.
  • Support and improve Snowflake deployment processes across environments so that configuration and platform changes are consistent, reviewable, and reliable.
  • Collaborate closely with application engineering, data engineering, security, and other teams; contribute to daily stand-up meetings and document operational procedures, known issues, and recovery steps.

Desired Skills and Experience

  • 3-5 years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, or production infrastructure operations, with hands-on responsibility for AWS environments.
  • Strong Linux administration and troubleshooting skills, including system services, networking, permissions, logs, and command-line diagnostics.
  • Hands-on working knowledge of AWS services used by the role, including Amazon S3, AWS Glue, Amazon Kinesis, AWS PrivateLink, security groups, and multi-account operational practices.
  • Good understanding of cloud networking fundamentals, including DNS, TCP/IP, routing, private endpoints, security groups, TLS, and troubleshooting application-to-service connectivity.
  • Experience supporting Snowflake connectivity and deployment processes, including diagnosing access and network issues and coordinating changes across multiple environments.
  • Experience with scripting and operational automation using tools such as Bash or Python; familiarity with CI/CD and Infrastructure as Code practices is highly desirable.
  • Experience with monitoring, logging, alerting, incident troubleshooting, and creating practical operational documentation and runbooks.
  • Well-organized and self-motivated, with strong multi-tasking skills and a positive, professional approach to working across technical and non-technical teams.

Similar jobs

Data Site Reliability Engineer (SRE)

General Dynamics Information TechnologyUnited States· 2 days ago
RemoteInformation Technology$111k–$150k/yrapply on gdit.wd5.myworkdayjobs.com