Jobs · Engineering · California

Senior MLOps Technical Lead

Shrive Technologies · Cupertino, CA · 5 days ago
EngineeringFull-time

About the role

We are looking for an experienced engineer to design and build the evaluation platform that our entire organization will rely on. This role starts with a blank page and requires a builder's mindset, comfort with ambiguity, and the ability to deliver robust, well-tested systems.

Responsibilities

  • Design and build evaluation systems for ML and LLM applications, including test harnesses, benchmark datasets, automated scoring (including LLM-as-judge approaches), and regression detection.
  • Develop reusable AI agent skills, plugins, and maintain internal marketplace infrastructure to extend and scale Data and AI/ML capabilities across the organization.
  • Embed automated quality gates into CI/CD pipelines (e.g., GitHub Actions, GitLab CI, Jenkins, Buildkite) to ensure reliability and performance.
  • Implement observability and monitoring for distributed systems, including LLM-specific tools (e.g., LangSmith, Langfuse, Arize Phoenix, Braintrust, W&B Weave).
  • Debug complex distributed systems under pressure, including production incident response, root-cause analysis, and blameless postmortems.
  • Translate evaluation results into clear findings and recommendations for both technical and non-technical stakeholders.
  • Work with agentic frameworks and orchestration patterns (e.g., multi-agent systems, tool use, RAG pipelines) and address their distinct failure modes.
  • Manage prompt workflows, model routing, or fine-tuning, and evaluate changes across model versions and providers.
  • Apply expertise in causal inference and measurement strategy, including causal graphs, ontologies, and knowledge graphs, to drive rigorous, decision-grade data analysis.

Requirements

  • 5+ years of experience in ML engineering, MLOps, platform engineering, or SRE, with 2+ years working hands-on with LLMs or LLM-powered applications in production.
  • Strong software engineering skills in Python (TypeScript is a plus), with a track record of building reliable, well-tested internal platforms and tooling.
  • Deep familiarity with CI/CD systems (e.g., GitHub Actions, GitLab CI, Jenkins, Buildkite).
  • Experience with observability and monitoring stacks (e.g., OpenTelemetry, Datadog, Grafana/Prometheus).
  • Experience with agentic frameworks, prompt management, model routing, or fine-tuning workflows.
  • Background in statistics or experimentation (A/B testing, significance testing, sampling strategies for human review).
  • Experience operating in regulated or high-stakes domains where agent errors carry real business or customer impact.

Skills

  • Python (TypeScript preferred)
  • CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, Buildkite)
  • Observability tools (OpenTelemetry, Datadog, Grafana/Prometheus, LangSmith, Langfuse, Arize Phoenix, Braintrust, W&B Weave)
  • Agentic frameworks and orchestration patterns (multi-agent systems, tool use, RAG pipelines)
  • Prompt management, model routing, and fine-tuning workflows
  • Causal inference, measurement strategy, and knowledge graphs
  • Debugging complex distributed systems and incident response
  • Cross-functional communication and stakeholder management

Similar jobs