Site Reliability Engineer (SRE), Data Products
CentralReach · Holmdel, NJ · Today
$135k–$160k/yrFull-time
Key Accountabilities
- Support deployment, configuration, and daily operations across several AWS accounts and the Linux servers and workloads running in those environments.
- Administer, monitor, and troubleshoot Amazon S3 buckets, AWS Glue jobs, Amazon Kinesis streams, AWS PrivateLink connectivity, security groups, and related AWS networking and access dependencies.
- Investigate and resolve infrastructure incidents, deployment failures, service interruptions, and connectivity issues, with a focus on restoring service quickly and preventing recurrence.
- Maintain and improve repeatable deployment and operational processes for AWS infrastructure and workloads through automation, version-controlled configuration, and clear runbooks.
- Support secure connectivity to Snowflake accounts, including troubleshooting network paths, private connectivity, endpoint configuration, and environment-specific access issues.
- Support and improve Snowflake deployment processes across environments so that configuration and platform changes are consistent, reviewable, and reliable.
- Collaborate closely with application engineering, data engineering, security, and other teams; contribute to daily stand-up meetings and document operational procedures, known issues, and recovery steps.
Desired Skills and Experience
- 3-5 years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, or production infrastructure operations, with hands-on responsibility for AWS environments.
- Strong Linux administration and troubleshooting skills, including system services, networking, permissions, logs, and command-line diagnostics.
- Hands-on working knowledge of AWS services used by the role, including Amazon S3, AWS Glue, Amazon Kinesis, AWS PrivateLink, security groups, and multi-account operational practices.
- Good understanding of cloud networking fundamentals, including DNS, TCP/IP, routing, private endpoints, security groups, TLS, and troubleshooting application-to-service connectivity.
- Experience supporting Snowflake connectivity and deployment processes, including diagnosing access and network issues and coordinating changes across multiple environments.
- Experience with scripting and operational automation using tools such as Bash or Python; familiarity with CI/CD and Infrastructure as Code practices is highly desirable.
- Experience with monitoring, logging, alerting, incident troubleshooting, and creating practical operational documentation and runbooks.
- Well-organized and self-motivated, with strong multi-tasking skills and a positive, professional approach to working across technical and non-technical teams.