Jobs · Engineering

Principal ML Ops Engineer

TEKTRND · United States · 1 mo ago
RemoteRemoteEngineeringContract

About the role

We are seeking an experienced ML Ops engineer to lead the architecture and implementation of scalable, production-grade AI inference solutions. You will work closely with our product and research teams to scale SOTA deep learning products and software, focusing on building and releasing high-performance AI runtimes.

Responsibilities

  • Architect and manage scalable model training and deployment pipelines for enterprise clients.
  • Lead the strategy for managing and releasing upstream and midstream AI product builds.
  • Design and implement automated testing frameworks to ensure model correctness, responsiveness, and efficiency.
  • Troubleshoot, debug, and upgrade mission-critical Dev & Test pipelines.
  • Define and deploy cybersecurity measures, including continuous vulnerability assessment and risk management for AI systems.
  • Collaborate with cross-functional teams to define market requirements and establish best practices for LLMOps.
  • Stay at the forefront of AI technologies and standards, driving innovation within our practice.

Qualifications

  • 5+ years of experience in ML Ops, DevOps, and Automation, with a focus on enterprise software deployment.
  • Expertise with Git, Github Actions, Terraform, Jenkins, Ansible, and modern automation/monitoring technologies.
  • Extensive experience administering Kubernetes/OpenShift in production environments.
  • Deep understanding of Agile development methodologies.
  • Proven experience with at least one major cloud provider: AWS, GCP, Azure, or IBM Cloud.
  • Expert-level Python programming skills.
  • Advanced troubleshooting and systems-thinking skills.
  • Experience contributing to open-source AI/ML projects (e.g., vLLM) is a strong plus.
  • Bachelor's degree or higher in Computer Science or a related discipline is preferred, but we prioritize practical experience and technical excellence.

Similar jobs