Site-Reliability Engineer (W2 Role) - Need Locals to AZ for In-person Interview
Saransh Inc · Scottsdale, AZ · 3 mo ago
On-siteEngineeringContract
About the role
The Site Reliability Engineer will be responsible for designing, implementing, and maintaining reliable, high-performance systems.
Responsibilities
- Design, implement, and maintain reliable, high-performance systems
- Write automation scripts and build dashboards for Application Performance Management
- Manage transaction journeys using programming languages such as Python, Ansible, and Node.js
- Implement Cloud observability using OTEL for real-time monitoring, distributed tracing, and incident resolution
- Manage containerized apps in GKE/RKE/AKE environments
- Work with specific GraphQL frameworks and troubleshoot networking protocols
- Manage application availability and improve gating and detection for a 24x7 high availability platform
- Manage in-memory caching solutions and implement tools like Redis
- Monitor and troubleshoot HashiCorp Vault environments
Requirements
- Experience in large-scale, high-performance applications in a hybrid environment
- Experience with Terraform, CI/CD (GitHub Actions), and Helm
- Knowledge of monitoring tools like Prometheus, Grafana, Splunk, App-dynamics, Dynatrace, and Dynatrace
- Experience with specific GraphQL frameworks and networking protocols
- Experience with in-memory caching solutions and Redis
- Experience with HashiCorp Vault and monitoring tools like Spanner and Firestore
- Experience with Enterprise-level infrastructure and operations
- Experience with High Availability and distributed systems, Linux, and Windows administration
- Experience with troubleshooting and support
Qualifications
- 7+ years of experience in service reliability/operations
- Mandatory skills: Google Cloud Platform, Kubernetes, Terraform, CI/CD, Helm, Python, Ansible, Node.js, Prometheus, Grafana, Linux, Redis, Clickhouse, postgres, Mongo, or any time-series databases
- Experience with specific GraphQL frameworks and networking protocols
- Experience with in-memory caching solutions and Redis
- Experience with HashiCorp Vault and monitoring tools like Spanner and Firestore
- Experience with Enterprise-level infrastructure and operations
- Experience with High Availability and distributed systems, Linux, and Windows administration, troubleshooting, and support
Skills
- Python
- Helm
- Google Cloud Platform
- Kubernetes
Benefits
N/A
Pay
N/A
Schedule
N/A