Lead AI/ML Engineer
Apex Systems · Dearborn, MI · 1 wk ago
EngineeringContract
About the role
Support the client's AI and ML engineering capability, including model fine-tuning oversight, agentic orchestration architecture, and LLM evaluation.
Responsibilities
- Oversee vendor fine-tuning of Google Cloud Vertex AI using proprietary diagnostic data, ensuring compliance with Ford's IP protection requirements and model weight storage architecture.
- Design and build the client's Orchestration Layer—the integration framework that connects external AI engine with other internal AI engines and platform services.
- Evaluate AI engine outputs against defined accuracy, latency, and first-time fix rate metrics; drive iterative improvement through structured feedback loops.
- Define model evaluation frameworks and acceptance criteria for AI-generated triage recommendations, ensuring clinical accuracy before dealer-facing deployment.
- Build internal tooling for model monitoring, drift detection, and retraining triggers within the client's GCP environment.
- Collaborate with the client's data engineering team to define data preparation and feature engineering requirements that support model fine-tuning and inference quality.
- Partner with the GCP Cloud Engineers to ensure model artifact storage, versioning, and access controls comply with the client's IP and security policies.
- Contribute to the long-term insourcing roadmap by documenting model architectures, training pipelines, and prompt frameworks in sufficient detail to enable internal replication.
- Represent AI and ML engineering in architecture reviews and vendor technical discussions.
Skills
Required
- Technical Communication – 2–5 years translating complex technical concepts (e.g., ML model behavior, data pipeline architecture, platform design decisions) into clear documentation, proposals, and presentations for technical and non-technical audiences, including engineering leads and product stakeholders.
- Communications – 2–5 years of demonstrated ability to communicate effectively across cross-functional teams, including facilitating technical discussions, contributing to design reviews, and keeping stakeholders aligned on project status, risks, and decisions.
- Google Cloud Platform – 2–5 years of hands-on experience with GCP services relevant to AI/ML and data workloads, including Vertex AI, BigQuery, GCS, Dataflow, or Cloud Composer, with the ability to deploy and manage workloads in a production cloud environment.
- TensorFlow – 2–5 years building, training, and evaluating machine learning models using TensorFlow or TensorFlow Extended (TFX), including experience with model versioning, pipeline integration, and deploying models to production serving infrastructure.
- Data Governance – 2–4 years applying data governance principles including data lineage, access controls, metadata management, and compliance standards to ensure telemetry and ML datasets meet quality, security, and regulatory requirements.
- Machine Learning – 3–5 years of applied ML experience including feature engineering, model selection, training, validation, and deployment. Comfortable working with both structured and unstructured data in the context of real-world engineering or automotive telemetry use cases.
- Python – 3–5 years writing production-quality Python for data engineering, ML pipeline development, or platform tooling. Proficiency with Pandas, NumPy, scikit-learn, and TensorFlow expected, along with familiarity with testing and version control.
- Artificial Intelligence & Expert Systems – 3–5 years of experience designing or working with AI systems, including large language models, expert systems, or intelligent automation within developer or data workflows. Understands model lifecycle management, prompt engineering, and responsible AI practices.
Preferred
- Telematics – 1–3 years of exposure to telematics data systems, including vehicle data collection, event streaming, or connected vehicle platforms. Familiarity with ingesting, processing, and applying telematics data to ML or analytics use cases is a strong plus.
Experience
- 5 or more years of professional experience in machine learning engineering, AI systems development, or applied AI research.
- Hands-on experience fine-tuning LLMs in a cloud environment, with specific preference for Google Cloud Vertex AI or equivalent managed ML platforms.
- Demonstrated experience building agentic AI systems using frameworks such as LangChain, LangGraph, Google Agent Builder, or equivalent orchestration tooling.
- Proficiency in Python and ML development tooling including Hugging Face, PyTorch or TensorFlow, and MLflow or Vertex AI Experiments.
- Experience designing and evaluating LLM outputs for production systems, including prompt engineering, retrieval-augmented generation (RAG) architectures, and model evaluation metrics.
- Strong understanding of MLOps practices including model versioning, deployment pipelines, monitoring, and retraining workflows on GCP.
- Experience working in regulated or IP-sensitive environments where model artifact ownership and data governance are active concerns.
- Strong written and verbal communication skills; ability to translate technical AI concepts for non-technical executive stakeholders.
Preferred
- Experience in automotive diagnostics, vehicle telematics, or connected vehicle platforms.
- Familiarity with Diagnostic Trouble Code (DTC) data, Over-the-Air (OTA) update systems, or repair order (RO) data structures.
- Experience with multi-agent AI systems and tool-use patterns in production.
- Google Cloud Professional Machine Learning Engineer certification.
Education
- Bachelor's Degree
Schedule
Hybrid / 4 days per week in the office.