Senior AI Site Reliability Engineer, AI.x
Roles & Responsibilities
Lead automation-first initiatives to eliminate toil and manual interventions, defining and executing the strategic roadmap for reliability, observability, and self-healing systems across AI.x platforms.
Design and implement robust CI/CD pipelines enabling one-touch deployments with automated testing, validation, and rollback capabilities to accelerate delivery velocity and reduce deployment risk.
Implement comprehensive observability frameworks for real-time monitoring of AI services, including metrics, logs, and traces, with intelligent alerting and automated diagnostics to minimize MTTD and MTTR.
Participate in on-call rotation providing 24/7 support for production AI systems, ensuring rapid incident response, root cause analysis, and resolution with measurable SLO targets.
Establish and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, and incident response runbooks to drive continuous reliability improvements.
Champion Infrastructure-as-Code (IaC) practices and automate environment provisioning, configuration management, and deployment processes to ensure consistency, repeatability, and operational efficiency.
Collaborate seamlessly with AI Engineering teams to integrate SRE practices early in the development lifecycle, promoting a culture of reliability and shared responsibility.
Proactively identify and resolve reliability, performance, and scalability issues through data-driven analysis, capacity planning, and system optimization.
Implement and maintain monitoring, alerting, and incident response frameworks to ensure system health and reliability, maximizing production availability.
Champion reliability, monitoring, observability, and operational best practices for AI systems and data pipelines, establishing patterns and standards for the organization.
Required Qualifications
- 8+ years of software engineering experience, with 4+ years as a hands-on Site Reliability Engineer in startups and/or large organizations.
- Bachelor’s degree in Computer Science or related field, or equivalent experience.
- 5+ years building complex products from scratch, running them in production, and ensuring operational reliability.
- 3+ years working with containers and cloud-native applications, operationalizing them in the public cloud with infrastructure as code and CI/CD pipelines.
- 3+ years of experience working in high-availability hybrid-cloud environments.
Preferred Qualifications
- Strong computer science fundamentals and experience across the tech stack.
- Experience with proprietary or open-source LLMs (e.g., Gemini, Claude, OpenAI), deploying LLM-powered applications to production and maintaining availability.
- Strong written and verbal communication skills to clearly convey ideas and feedback.
- Strong understanding of observability, incident management and reliability engineering principles.
- Mindset of continuous learning and improvement, adept at both giving and receiving feedback.
- Able to troubleshoot complex problems with ambiguous or incomplete data in distributed systems.
- Curiosity about new technologies and processes, proactively sharing knowledge and seeking improvement.
- Experience with Terraform and Google Cloud Platform.