Jobs · Engineering

Senior Systems Reliability Engineer II

ThoughtSpot · Mountain View, CA · 1 mo ago
RemoteRemoteEngineeringFull-time

About the role

As part of the ThoughtSpot SRE team, you will be on the cutting edge of operational intelligence. You will ensure service reliability and act as a trusted partner for our customers, leveraging AI/ML to deliver timely updates, meaningful solutions, and predictive improvements. You are the bridge between our customers and engineering, combining deep systems expertise with a genuine passion for customer success.

Responsibilities

  • Act as the primary point of contact for customer-facing technical issues related to our SaaS platform, including data connectivity, report errors, performance concerns, access problems, data inconsistencies, software bugs, and integration challenges.
  • Understand and empathize with the challenges ThoughtSpot users face, offering tailored solutions to improve their experience.
  • Provide timely, accurate, and clear updates to customers, consistently meeting SLAs and driving issues through to full resolution via tickets and calls.
  • Create and maintain knowledge-base articles to empower customer self-service and improve support efficiency.
  • Maintain, monitor, and troubleshoot ThoughtSpot cloud infrastructure using tools like Grafana, Prometheus, Datadog, and Splunk.
  • Maintain system health and performance through metrics, logs, and dashboards to detect and prevent issues proactively.
  • Implement and leverage AI/ML-driven solutions for proactive observability, predictive anomaly detection, and intelligent alerting to enhance service reliability and reduce Mean Time to Resolution (MTTR).
  • Develop and implement automation and best practices to streamline operations and strengthen system reliability.
  • Optimize SRE workflows with AI tools to boost operational effectiveness.
  • Participate in on-call rotations, lead incident reviews, and conduct thorough root cause analyses to drive continuous improvement.
  • Work cross-functionally with Engineering to define and implement tools that enhance debuggability, supportability, availability, scalability, and performance.

Requirements

  • B.S. in Computer Science or equivalent relevant experience.
  • Proven experience troubleshooting complex Linux systems and managing virtualization and cloud platforms (VMware, AWS, Azure, GCP).
  • Hands-on experience with monitoring tools such as Grafana, Prometheus, Datadog, or Splunk.
  • Demonstrated experience and a keen interest in leveraging AI/ML principles to address SRE challenges — including AIOps, predictive maintenance, and intelligent automation.
  • Prior experience in enterprise customer support, including on-call rotations and incident management, with the ability to lead root cause analyses.
  • Strong problem-solving and algorithmic thinking with a solid understanding of system internals.
  • Excellent verbal and written communication skills with the ability to work independently and cross-functionally in fast-paced environments.
  • Familiarity with scripting and programming languages such as Python, Go, Bash, or Java.
  • Exposure to infrastructure and service monitoring frameworks with the ability to analyze data to ensure high availability.

Qualifications

  • Good to Have:
  • Experience partnering with Engineering to design and implement mission-critical tooling and automation that advances system debuggability, high availability, elastic scalability, and performance.
  • Experience with alerting strategies and monitoring system tuning to minimize alert fatigue and optimize Mean Time to Acknowledge (MTTA).
  • Familiarity with C/C++ or other low-level systems languages.

Skills

  • Comfortable and confident integration of artificial intelligence into daily workflow to increase productivity and quality.
  • Hands-on experience to leverage AI tools (industry-leading LLMs) to increase productivity, automate routine tasks, and improve work quality.
  • Write effective prompts to get the most accurate and creative results from AI tools.

Benefits

  • Competitive salary and benefits package.
  • Opportunities for professional growth and career advancement.
  • A collaborative work environment where your input and expertise directly impact customer experience and platform reliability.

Pay

  • Details TBD.

Schedule

  • Hybrid Work at ThoughtSpot Spotters are expected in-office 3 days per week to experience the energy of their local office.

What We Offer

  • ThoughtSpot is the Agentic Analytics Platform that empowers every enterprise to transform insights into action, on a mission to make the world more fact driven.
  • We hire people with unique identities, backgrounds, and perspectives - this balance-for-the-better philosophy is key to our success.
  • We welcome different backgrounds, identities, and experiences, and we work to create a place where everyone can be themselves and do their best work.

Mandatory And Required Skills For All ThoughtSpot Roles

  • Spotters are expected to demonstrate AI literacy and workflow integration to include the ability to: Comfortably and confidently integrate artificial intelligence into their daily workflow to increase productivity and quality.
  • Hands-on experience to leverage AI tools (industry-leading LLMs) to increase productivity, automate routine tasks, and improve work quality.
  • Speak to the experience of using AI for research, content creation, and document summarization while maintaining ownership of judgment and final decisions.
  • Write effective prompts to get the most accurate and creative results from AI tools.

Ai Mindset For All Spotters At ThoughtSpot

  • Curiosity in exploring new AI tools.
  • Adaptability to quickly learn and implement new, emerging AI technologies.
  • Critical thinking to know when to identify when AI should be used versus when human judgement is necessary.

ThoughtSpot For All

  • ThoughtSpot is the Agentic Analytics Platform that empowers every enterprise to transform insights into action, on a mission to make the world more fact driven.
  • We hire people with unique identities, backgrounds, and perspectives - this balance-for-the-better philosophy is key to our success.
  • We welcome different backgrounds, identities, and experiences, and we work to create a place where everyone can be themselves and do their best work.

Similar jobs