Jobs · Information Technology

(US) Principal ML System Engineer

PointClickCare · United States · 3 days ago
RemoteRemoteInformation TechnologyFull-time

This team serves as the product owner for the machine learning platform capabilities within PointClickCare, working closely with other engineering teams to identify, build, and support traditional machine learning (ML) and hybrid ML/LLM solutions. This centralized team with deep specialization integrates closely with key horizontal partners to ensure delivery of safe, scalable, and high-impact AI products.

Responsibilities

  • Partner with product and engineering leadership to translate business and product objectives into a multi-quarter technical strategy and roadmap for the ML platform.
  • Define the reference architectures and standards for scalable data and ML pipelines spanning model training, evaluation, deployment, and serving that engineering teams across the organization build upon.
  • Set the direction and best practices for MLOps across the company — including CI/CD for models, model registry, feature stores, and experiment tracking — and drive build-vs-buy decisions for core platform components.
  • Establish the practices and architecture for reliability, observability, and performance of ML systems in production, including monitoring, alerting, and automated remediation.
  • Establish the security architecture for the ML platform, including authentication, role-based access control, audit logging, and compliance monitoring, and ensure adoption across teams.
  • Define secure, cost-efficient integration and infrastructure patterns for connecting the platform with existing systems, APIs, and data sources at scale.
  • Provide technical leadership and mentorship across engineering teams, guiding senior engineers and influencing the org-wide technical roadmap for ML infrastructure.

Qualifications & Skills

  • Expert level in Python and Java with strong software engineering fundamentals.
  • Deep experience designing and building ML platforms and ML Ops workflows at scale, familiarity with tools such as MLFlow, Kubeflow, Ray, and model-serving frameworks or equivalents.
  • Extensive experience with cloud platforms (AWS, Azure, and/or GCP), containerization, and orchestration (Docker, Kubernetes).
  • Demonstrated track record of setting technical direction and driving org-wide technical initiatives across multiple teams.
  • Preferred Bachelor’s degree or higher in Computer Science, Machine Learning, or a related field.
  • Sufficient familiarity with Azure Machine Learning components, Databricks processing and serverless environments, and ML Frameworks to support strategic decision making.
  • Demonstrable history of leading and sustaining build out of critical cross-team systems.
  • Experience implementing security at scale including role-based access control, multi-factor authentication, network security best practices, and compliance monitoring.
  • Experience optimizing large model training and inference (including LLM serving) for performance and cost.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications or analyzing resumes.

Pay

US base salary range for this position: $195,000–$217,000 (not overtime eligible) plus bonus and benefits. Compensation is assessed individually and aligned to experience, skills, and market context.

Similar jobs

Principal ML Engineer

Prodege, LLCUnited States· 1 mo ago
RemoteEngineering$300k–$375k/yrapply on prodege.wd108.myworkdayjobs.com

ML systems engineer

MyndstackUnited States· 2 wk ago
RemoteInformation Technologyapply on myndstack.io

ML Systems Engineer

EM DASH LABSTexas, United States· 1 wk ago
Information Technology$14k–$16k/moapply on emdashlabs.ai