Site Reliability Engineer (SRE) – II
EngineeringFull-time
Incident Management
- Handle complex incidents and ensure service uptime.
- Lead troubleshooting efforts for high-impact production issues, providing detailed root cause analysis (RCA) and preventative measures.
- Participate in on-call rotations, acting as an escalation point for Level 1 SREs during major incidents.
Automation & Infrastructure as Code (IaC)
- Develop and maintain automation scripts and infrastructure using tools like Terraform, Ansible, or CloudFormation.
- Implement automation solutions to eliminate manual tasks and improve system reliability, scalability, and performance.
Performance & Scalability
- Analyze system performance and recommend optimizations for scalability and reliability.
- Support capacity planning efforts by monitoring system metrics, traffic patterns, and usage trends to predict future resource needs.
System Design & Architecture
- Collaborate with software engineering teams to influence the design of new services and applications, ensuring they are scalable, reliable, and resilient from the start.
- Contribute to architectural decisions, ensuring alignment with best practices in fault tolerance, redundancy, and recovery.
Monitoring & Observability
- Build and maintain robust monitoring, alerting, and observability solutions to proactively detect and resolve issues before they impact end users.
- Optimize existing monitoring tools (e.g., Prometheus, Grafana, Datadog, Dynatrace) and build custom dashboards for better visibility into system health.
Security & Compliance
- Ensure systems and infrastructure are secure, compliant, and aligned with organizational policies and industry best practices.
- Assist with vulnerability management, system patching, and implementing security measures to protect the integrity and availability of services.
Continuous Improvement
- Lead efforts to continuously improve operational processes, tools, and workflows.
- Implement and enforce best practices in deployment, monitoring, and incident management to improve overall system reliability and reduce downtime.
Basic Qualification
- Minimum 5 years of experience in site reliability engineering, DevOps, systems administration, or related roles.
- Strong experience with Linux/Unix administration and proficiency in scripting (e.g., Python, Bash, Go).
- Deep understanding of cloud platforms (AWS, GCP, Azure) and related services (EC2, S3, Lambda, Kubernetes, etc.).
- Experience with containerization and orchestration technologies like Docker and Kubernetes.
- Proficiency with monitoring and observability tools such as dynatrace, Prometheus, Grafana, Datadog, ELK Stack, or similar platforms.
- Strong understanding of networking fundamentals (DNS, HTTP, TCP/IP), load balancing, and CDNs.
- Experience with CI/CD tools (Jenkins, GitLab CI, CircleCI) and infrastructure automation (Terraform, Ansible, Puppet).
- Familiarity with distributed systems and microservices architecture.
- Excellent problem-solving and troubleshooting skills, especially in diagnosing production issues in high-scale environments.
Preferred
- Background in MLOps, data engineering, and/or cloud-native AI deployment.
- Strong communication and documentation abilities.
- Knowledge of security best practices for AI and cloud infrastructure.
- Contributions to open source AI/SRE projects or relevant technical communities.
- Proven track record of managing complex infrastructure, troubleshooting production issues, and optimizing system performance.