Model Operations Engineer
MANTECH · Ashburn, VA · 4 days ago
On-siteInformation TechnologyFull-time
Responsibilities
- Lead the integration and deployment of trained AI/ML models into production environments (e.g., cloud, edge devices) using ML Ops best practices.
- Develop and optimize model training & inference pipelines for real-time execution and efficiently handle large-scale data processing.
- Work with data science teams to structure automated ML model health monitoring and refresh capabilities.
- Implement continuous integration, delivery and training (CI/CD/CT) workflows with commercial and open-source modeling platforms/services.
- Collaborate with cross-functional teams (e.g., Software Engineering, Data Science) to integrate and test multiple candidate AI/ML models and applications for operational assessment.
Requirements
- Expertise with ML Ops tools and frameworks such as Mlflow, Kubeflow, Airflow and implementing monitoring/drift detection capabilities (e.g. Alibi, Grafana).
- Experience automating workflow orchestration to handle both batch and real-time streaming data processing for model inference.
- Hands-on experience productionizing models, including experience optimizing for inference speed, containerization (e.g., Docker), and with multi-cloud deployment platforms (e.g., AWS, Azure, GCP).
- Proficiency in Python, Scala and Java with strong understanding of high-performance computing and GPU acceleration.
- Experience with data engineering Extract, Transform and Load (ETL) workflows across various relational/non-relational databases (Oracle/Postgres, MongoDB) and cloud endpoint services, e.g. (Lambda, GraphQL etc.).
- Experience in using deep learning frameworks (PyTorch, TensorFlow, Keras) and computer vision libraries (OpenCV, SimpleITK, ITKm VTK).
- Experience with biometric or image recognition algorithms and associated predictive analytics pipelines.
- Experience with GPU-based infrastructure and performance optimization.
Qualifications
- HS Diploma/GED and 15-20 years of experience, AS/AA and 13-18 years, BS/BA and 7-12 years or MS/MA/MBA and 5-9 years or PhD/Doctorate and 3-7 years.
Skills
- Deep expertise and experience with predictive modeling lifecycles.
- Hands-on experience with machine learning tools and frameworks.
- A pragmatic, customer-centric approach to applying ML models to solve complex problems.
Benefits
- This is currently a hybrid position with two days onsite in Ashburn, VA and three days remote.
Pay
- Commensurate with experience.
Schedule
- Hybrid: Two days onsite in Ashburn, VA and three days remote.