Overwatch - Observability and Evaluation Engineer
AMroute LLC · Bluffdale, UT · 1 wk ago
On-siteOTHRFull-time
Location: Charlotte, North Carolina
Duration: 12+ months
Schedule: 40 hours per week, 100% onsite
Visa: Any, as long as eligible to work (local candidates preferred)
About the role
We are seeking an Overwatch Observability and Evaluation Engineer to implement telemetry, tracing, dashboards, evaluation suites, service objectives, runbooks, and release-readiness evidence for high-priority AI agent releases. This role is focused on ensuring agentic AI systems are measurable, reliable, production-ready, and continuously monitored.
Responsibilities
- Design and implement observability frameworks for LLM applications and AI agent workflows.
- Build telemetry, tracing, logging, metrics, and dashboarding capabilities for production AI systems.
- Develop evaluation suites to measure prompt quality, model behavior, task success, groundedness, latency, reliability, and safety.
- Define and track SLOs, SLIs, service readiness criteria, and operational health indicators.
- Create automated test and evaluation workflows for model, prompt, and agent performance analysis.
- Build runbooks and operational readiness evidence for production releases.
- Partner with AI engineering, platform, DevOps, product, and risk teams to support priority agent launches.
- Analyze production signals, identify performance issues, and recommend improvements for reliability and user outcomes.
Requirements
- Bachelor's degree in Computer Science, Data Science, AI/ML, Engineering, or a related field.
- Experience implementing observability, evaluation, or monitoring for AI/ML, LLM, or agent-based systems.
- Strong Python scripting and automation skills.
- Ability to support production-grade AI releases with clear documentation, metrics, and operational readiness discipline.
Skills
Must-Have
- LLM Evaluation
- AI Agent Evaluation
- Observability
- Tracing
- Telemetry
- Metrics
- Dashboards
- SLOs
- Runbooks
- Test Automation
- Prompt Performance Analysis
- Model Performance Analysis
- Python
- Production Operations
Preferred
- LangSmith
- OpenTelemetry
- Grafana
- Prometheus
- Datadog
- CloudWatch
- Azure Monitor
- AI Reliability Engineering
- LLMOps
- MLOps
- Incident Response
- Release Readiness