LLM Ops Engineer
Litera · Denver, CO · 1 mo ago
Hybrid$105k–$130k/yrFull-time
About the role
This is a hybrid role based in Denver, CO with the expectation to be in the office at least 3 days a week for collaboration and connection.
Responsibilities
- Build and operate a scalable AI platform that enables engineering teams to seamlessly access and deploy models across multiple providers and environments.
- Ensure high availability and resiliency of AI services through intelligent routing, failover strategies, and production-grade infrastructure.
- Establish secure and compliant AI operations by protecting model access, safeguarding sensitive data, and enforcing governance standards.
- Create a consistent developer experience through unified APIs, self-service capabilities, tooling, and best practices that accelerate AI adoption.
- Optimize AI platform performance, reliability, and cost efficiency through proactive monitoring, analytics, and provider strategy management.
- Lead the evolution of Litera’s AI operations capabilities by evaluating emerging technologies and recommending scalable solutions.
- Deliver observability and operational excellence through dashboards, alerting, quality monitoring, and service-level metrics.
- Support the safe deployment of AI solutions by implementing testing frameworks, quality controls, and production readiness standards.
Requirements
- 3+ years of experience in DevOps, Platform Engineering, MLOps, or a related field, including hands-on experience operating LLMs in production environments.
- Experience deploying, managing, and scaling models across multiple AI providers such as OpenAI, Anthropic, Azure OpenAI, AWS Bedrock, or Google Vertex AI.
- Strong expertise in building highly available, secure infrastructure, including load balancing, failover strategies, secrets management, and access controls.
- Experience with API management, gateway technologies, and production-grade AI service operations.
- Strong Python programming skills with experience developing and supporting scalable systems.
- Proven ability to solve complex technical challenges and thrive in a fast-paced, evolving environment while collaborating across teams.
Qualifications
- Experience fine-tuning or training large language models for domain-specific applications.
- Familiarity with ML orchestration tools and frameworks such as Kubeflow, MLflow, or Apache Airflow.
- Experience with LLM evaluation frameworks, retrieval-augmented generation (RAG), vector databases, or inference optimization techniques.
- Knowledge of infrastructure-as-code, Kubernetes, compliance frameworks, or large-scale AI cost optimization strategies.
Skills
- Python programming skills.
- Experience with API management and gateway technologies.
- Hands-on experience operating large language models in production environments.
- Experience with ML orchestration tools and frameworks.
- Experience with LLM evaluation frameworks, RAG, vector databases, or inference optimization techniques.
- Knowledge of infrastructure-as-code, Kubernetes, compliance frameworks, or large-scale AI cost optimization strategies.
Benefits
Comprehensive benefits package including medical, dental, and vision coverage, a 401(k) with company match, and incentive and recognition programs.
Pay
The base salary range for this role is $105,000 to $130,000 USD. Final compensation will be determined based on experience, skills, education, and other relevant qualifications.
Schedule
This is a hybrid role based in Denver, CO with the expectation to be in the office at least 3 days a week for collaboration and connection.