Jobs · Information Technology · California

Manager, Machine Learning Engineering

GoFundMe · San Francisco, CA · 2 days ago
HybridInformation Technology$219k–$329k/yrFull-time

About the role

We’re looking for a Manager, Machine Learning Engineering (ML and AI Operations) to join our team. This role is located in the San Francisco Bay Area and requires an in-office presence 3 days a week.

Responsibilities

  • Own the reliability, scalability, and operational health of ML/AI production systems across GoFundMe, including training pipelines, feature stores, model serving, and monitoring/observability infrastructure.
  • Lead, hire, and grow a team of ML/AI operations engineers, setting technical direction through design reviews, architecture decisions, and shared best practices for production ML and AI systems.
  • Partner with data science and ML engineering teams to streamline the path from model development to production deployment, including CI/CD for ML, model packaging, versioning, and rollback strategies.
  • Establish ML operational excellence org-wide by driving standards for model observability (latency, errors, drift, calibration, business KPI deltas), automated retraining triggers, and incident response playbooks.
  • Build and mature on-call processes, SLOs/SLAs, and postmortem practices for ML/AI systems, treating model incidents with the same discipline as production infrastructure incidents.
  • Drive operational strategy for GoFundMe's generative AI systems alongside traditional ML, balancing innovation velocity with safety, compliance, cost, and reliability.
  • Collaborate cross-functionally with Product, Engineering, Design, and Legal/Privacy stakeholders to translate business goals into team priorities and measurable operational outcomes.
  • Manage vendor and platform relationships (e.g., cloud ML platforms, LLM providers) and make build-vs-buy calls that balance cost, control, and speed.
  • Report on team health, system reliability metrics, and operational risk to senior engineering leadership.

Requirements

  • 7+ years of hands-on experience building and shipping production machine learning systems, with demonstrated ownership of backend services and ML pipelines in a high-availability environment.
  • 1-3+ years of experience directly managing engineers, ideally in an MLOps, ML platform, or infrastructure context, with a track record of hiring and developing strong teams.
  • Strong proficiency in Python and ML libraries/frameworks such as PyTorch, TensorFlow, Scikit-learn, plus strong software engineering fundamentals (testing, code review, CI/CD, API design, performance, and reliability) — enough depth to stay hands-on and credible with your team.
  • Experience designing and operating real-time model serving at scale, including containerization, scalable inference, feature retrieval, and safe rollout strategies (canaries, shadowing, backward-compatible schema evolution).
  • Strong data engineering fluency: building reliable datasets and features using SQL, Spark/Databricks, and warehouse technologies (e.g., Snowflake), with an understanding of event semantics, identity resolution, and data quality controls.
  • Proven experience implementing ML monitoring for both technical and business metrics (drift, calibration, segment performance, latency, error budgets) and running models reliably in production.
  • Familiarity with generative AI/LLM infrastructure and operational considerations (latency, cost, safety guardrails) is a strong plus.
  • Able to break down ambiguous, high-impact problems, define crisp interfaces and success metrics, and deliver iteratively while managing stakeholder expectations across engineering leadership, product, and data science.
  • Strong leadership and mentoring skills and a proven ability to raise the bar on architecture, engineering quality, and operational rigor for production ML/AI systems.
  • Advanced degree (Master's or Ph.D.) in Computer Science, Statistics, Data Science, or a related technical field is preferred.

Qualifications

  • Passionate about using technology to make a positive impact.
  • Ability to work independently and collaboratively with a diverse team.
  • Strong communication and interpersonal skills.
  • Ability to manage multiple projects and priorities simultaneously.

Skills

  • Python
  • Machine Learning Libraries/Frameworks (PyTorch, TensorFlow, Scikit-learn)
  • Data Engineering (SQL, Spark/Databricks, Snowflake)
  • Model Serving (containerization, scalable inference, feature retrieval)
  • ML Monitoring (drift, calibration, segment performance, latency, error budgets)
  • Generative AI/LLM Infrastructure (latency, cost, safety guardrails)

Benefits

GoFundMe offers a competitive salary range of $219,000 - $329,000 annually, with additional benefits including healthcare, dental, vision, life insurance, and a 401(k) savings program. Geolocation differentials may apply depending on the work location. The company also provides equity and other benefits, as well as financial assistance for things like hybrid work, family planning, and generous parental leave, flexible time-off policies, and mental health and wellness resources to support your overall well-being.

Pay

The annual U.S. salary range for this full-time position is $219,000 - $329,000.

Schedule

This role requires an in-office presence 3 days a week.

Similar jobs