Jobs · Engineering · New Jersey

Machine Learning Platform - Associate - New York

Goldman Sachs · Jersey City, NJ · 1 mo ago
EngineeringFull-time

About the role

The successful candidate will be part of an expert team building and operating production-grade platform and backend systems leveraged by ML engineers and application teams across the entire firm. They will focus on enabling reliable, scalable, and observable deployment of Machine Learning and Large Language Models (LLMs).

Responsibilities

  • Deliver scalable, efficient, secure and automated processes for building, deploying and monitoring Machine Learning models
  • Enable solutions that provide business customers with the ability to leverage the latest and greatest AI/ML infrastructure, frameworks, and tooling to deliver high impact outcomes
  • Develop and demonstrate deep subject matter expertise on how to optimize machine learning model deployments to scale to the specific needs of each business customer
  • Author and maintain high quality documentation for both the engineering team as well as for business customers
  • Participate in on-call and support rotations, helping diagnose and resolve production issues
  • Continuously expand knowledge of platform architecture with a goal to take ownership of individual components
  • Stay up to date with advancements in AI/ML frameworks, model serving technologies, and GenAI infrastructure

Requirements

  • 2 years of experience in software engineering (backend, platform, or infrastructure)
  • 2 years of experience in Python or a similar backend programming language
  • 1 year of experience supporting production ML systems (MLOps, platform or inference-related work)
  • Basic understanding of APIs (REST or similar) and service-to-service communication
  • Experience working with containers (e.g., Docker)
  • Familiarity with Unix-based systems
  • Exposure to public cloud environments (e.g., AWS or GCP), including core concepts such as compute, storage, and basic IAM
  • Experience working with databases (SQL or NoSQL)
  • Solid grasp of software engineering fundamentals, including debugging, testing, and maintainable code design
  • Strong problem-solving skills and the ability to work effectively in a fast-paced, collaborative environment
  • Curiosity and a strong desire to keep learning—especially in the model inference and LLM platform space

Qualifications

  • 4 years of experience in software engineering (backend, platform, or infrastructure)
  • 4 years of experience supporting production ML systems (MLOps, platform or inference-related work)
  • 4 years of experience in Python or a similar backend programming language
  • Strong understanding of the end-to-end Model Development Lifecycle (MDLC)
  • Basic understanding of distributed systems concepts and exposure to observability concepts (logging, metrics, tracing)
  • Experience building containerized runtime environments for model serving (e.g. vLLM, SGLang, TensorRT, Triton, AWS Multi Model Server)
  • Experience with infrastructure-as-code tools, such as Terraform or CloudFormation
  • Experience with Kubernetes and other container orchestration platforms in the public cloud (e.g. AWS, GCP)
  • Experience building Machine Learning models with frameworks such as PyTorch and TensorFlow

Skills

  • Excellent communication skills and the ability to articulate complex technical concepts to both technical and non-technical stakeholders

Benefits

  • Commensurate with experience

Pay

  • TBD

Schedule

  • TBD

Similar jobs