Senior MLOps Technical Lead
Shrive Technologies · Cupertino, CA · 5 days ago
EngineeringFull-time
About the role
We are looking for an experienced engineer to design and build the evaluation platform that our entire organization will rely on. This role starts with a blank page and requires a builder's mindset, comfort with ambiguity, and the ability to deliver robust, well-tested systems.
Responsibilities
- Design and build evaluation systems for ML and LLM applications, including test harnesses, benchmark datasets, automated scoring (including LLM-as-judge approaches), and regression detection.
- Develop reusable AI agent skills, plugins, and maintain internal marketplace infrastructure to extend and scale Data and AI/ML capabilities across the organization.
- Embed automated quality gates into CI/CD pipelines (e.g., GitHub Actions, GitLab CI, Jenkins, Buildkite) to ensure reliability and performance.
- Implement observability and monitoring for distributed systems, including LLM-specific tools (e.g., LangSmith, Langfuse, Arize Phoenix, Braintrust, W&B Weave).
- Debug complex distributed systems under pressure, including production incident response, root-cause analysis, and blameless postmortems.
- Translate evaluation results into clear findings and recommendations for both technical and non-technical stakeholders.
- Work with agentic frameworks and orchestration patterns (e.g., multi-agent systems, tool use, RAG pipelines) and address their distinct failure modes.
- Manage prompt workflows, model routing, or fine-tuning, and evaluate changes across model versions and providers.
- Apply expertise in causal inference and measurement strategy, including causal graphs, ontologies, and knowledge graphs, to drive rigorous, decision-grade data analysis.
Requirements
- 5+ years of experience in ML engineering, MLOps, platform engineering, or SRE, with 2+ years working hands-on with LLMs or LLM-powered applications in production.
- Strong software engineering skills in Python (TypeScript is a plus), with a track record of building reliable, well-tested internal platforms and tooling.
- Deep familiarity with CI/CD systems (e.g., GitHub Actions, GitLab CI, Jenkins, Buildkite).
- Experience with observability and monitoring stacks (e.g., OpenTelemetry, Datadog, Grafana/Prometheus).
- Experience with agentic frameworks, prompt management, model routing, or fine-tuning workflows.
- Background in statistics or experimentation (A/B testing, significance testing, sampling strategies for human review).
- Experience operating in regulated or high-stakes domains where agent errors carry real business or customer impact.
Skills
- Python (TypeScript preferred)
- CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, Buildkite)
- Observability tools (OpenTelemetry, Datadog, Grafana/Prometheus, LangSmith, Langfuse, Arize Phoenix, Braintrust, W&B Weave)
- Agentic frameworks and orchestration patterns (multi-agent systems, tool use, RAG pipelines)
- Prompt management, model routing, and fine-tuning workflows
- Causal inference, measurement strategy, and knowledge graphs
- Debugging complex distributed systems and incident response
- Cross-functional communication and stakeholder management