Jobs · Engineering · New York

Machine Learning Engineer, Infra, AI for Drug Discovery

Genentech · New York, NY · Yesterday
Engineering$148k–$274k/yrFull-time

About the role

A healthier future drives us to innovate, advancing science and ensuring access to healthcare for generations. Roche’s Research and Early Development organisations (gRED and pRED) leverage AI, data, and computational sciences to accelerate drug discovery and development. The Computational Sciences Center of Excellence (CoE) unifies these efforts, harnessing data and AI to deliver transformative medicines.

Within Roche’s AI for Drug Discovery (AI4DD) group (Prescient Design), we build machine learning platforms that enable researchers and engineers to move models from experimentation into reliable scientific and production workflows. As a Machine Learning Infrastructure Engineer, you will help build and operate platforms supporting model deployment, evaluation, promotion, monitoring, and lifecycle management across the organization. This role focuses on our model-serving platform and broader infrastructure to simplify deployment, scaling, observability, and integration into scientific and agentic workflows.

Responsibilities

  • Design, implement, ship, and operate scalable model-serving infrastructure for machine learning, scientific, LLM, and agentic workloads.
  • Evolve our internal model deployment platform into a reliable, self-service platform for teams across the organization.
  • Improve platform scalability and reliability, including scale-to-zero, faster model startup, workload isolation, traffic management, and reduction of request failures and latency bottlenecks.
  • Build observability and operational tooling for model usage, latency, reliability, resource consumption, inference cost, bottlenecks, and service-level indicators.
  • Enhance model deployment usability through validated configuration interfaces, reusable deployment patterns, APIs, command-line tools, and documentation.
  • Converge real-time and batch inference workflows onto shared platform capabilities where appropriate.
  • Contribute to model lifecycle management infrastructure, including model registration and versioning, evaluation, promotion and release gates, monitoring, environment progression, and rollback.
  • Build event-driven integrations connecting model publication, evaluation, promotion, deployment, and retraining workflows.
  • Develop consistent metrics and evaluation signals for understanding model cost, quality, reliability, and fitness for downstream workflows.
  • Partner with machine learning, data, scientific, and platform teams to translate requirements into maintainable solutions and remove infrastructure bottlenecks.
  • Own workstreams from design through implementation and production support, using strong software-engineering practices including testing, reviews, documentation, and incremental delivery.

Requirements

  • BS or MS in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
  • 3+ years of relevant industry experience in software engineering, infrastructure engineering, platform engineering, DevOps, MLOps, or a related area.
  • Strong Python programming skills and experience building and shipping maintainable production software, services, automation, or developer tooling.
  • Demonstrated interest in hands-on implementation and production software delivery.
  • Experience designing, deploying, or operating cloud systems (preferably on AWS) using services such as EKS, EC2, S3, IAM, SQS, SNS, and CloudWatch.
  • Experience with containers, Kubernetes, Helm, and IaC tools such as Terraform or Pulumi.
  • Experience with CI/CD, Git-based development workflows, automated testing, and software release practices.
  • Ability to troubleshoot complex systems using metrics, logs, traces, events, and observability tools such as Datadog, Prometheus, Grafana, or OpenTelemetry.
  • Understanding of distributed-systems concepts such as concurrency, queuing, retries, timeouts, idempotency, backpressure, and failure recovery.
  • Ability to gather requirements, communicate technical tradeoffs, and document systems for users and engineers with varied infrastructure experience.
  • Demonstrated ability to independently deliver practical, incremental solutions while considering immediate needs and longer-term platform direction.

Preferred Qualifications

  • Familiarity with model-serving or workflow-orchestration frameworks such as KServe, Triton, vLLM, Ray Serve, Prefect, or Dagster.
  • Experience optimizing model startup time, request throughput, batching, autoscaling, or GPU utilization.
  • Familiarity with model registries, experiment tracking, model evaluation, promotion workflows, or MLOps platforms.
  • Experience building event-driven systems using queues, event buses, or workflow orchestrators.
  • Familiarity with online and offline model evaluation, model-quality monitoring, data drift, or regression analysis.
  • Experience supporting scientific computing, high-performance computing, distributed training, or large-scale data processing.
  • Strong interest in the life sciences and drug discovery.

Pay

The expected salary range for this position based on the primary location of California is $147,600–$274,000, and for New York, $141,100–$262,100. Actual pay will be determined based on experience, qualifications, geographic location, and other job-related factors permitted by law. A discretionary annual bonus may be available based on individual and Company performance.

Benefits

This position qualifies for the benefits detailed at the provided link.

Similar jobs