Jobs · Engineering · California

Staff Software Engineer - AI Platform

LinkedIn · Mountain View, CA · 1 wk ago
HybridEngineering$175k–$287k/yrFull-time

About the role

At LinkedIn, our approach to flexible work is centered on trust and optimized for culture, connection, clarity, and the evolving needs of our business. The work location of this role is hybrid, meaning it will be performed both from home and from a LinkedIn office on select days, as determined by the business needs of the team.

Join us to push the boundaries of scaling large models together. The team is responsible for scaling LinkedIn’s AI model training, feature engineering, and serving with hundreds of billions of parameters models and large-scale feature engineering infrastructure for all AI use cases—from recommendation models and large language models to computer vision models. We optimize performance across algorithms, AI frameworks, data infrastructure, compute software, and hardware to harness the power of our GPU fleet with thousands of the latest GPU cards. The team also works closely with the open-source community and has many open-source committers (TensorFlow, Horovod, Ray, vLLM, Huggingface, DeepSpeed, etc.). Additionally, this team focuses on technologies like LLMs, GNNs, incremental learning, online learning, and serving performance optimizations across billions of user queries.

Responsibilities

  • Owning the technical strategy for broad or complex requirements with insightful and forward-looking approaches that go beyond the direct team and solve large open-ended problems.
  • Designing, implementing, and optimizing the performance of large-scale distributed serving or training for personalized recommendation as well as large language models.
  • Improving the observability and understandability of various systems with a focus on improving developer productivity and system sustenance.
  • Mentoring other engineers, defining our challenging technical culture, and helping to build a fast-growing team.
  • Working closely with the open-source community to participate in and influence cutting-edge open-source projects (e.g., vLLMs, PyTorch, GNNs, DeepSpeed, Huggingface, etc.).
  • Functioning as the tech-lead for several concurrent key initiatives in AI Infrastructure and defining the future of AI Platforms.

Model Training Infrastructure

  • Designing and implementing high-performance data I/O.
  • Working with open-source technologies to identify and resolve issues in popular libraries like PyTorch, Huggingface, etc.
  • Enabling distributed training over hundreds of billions of parameter models.
  • Debugging and optimizing deep learning training.
  • Providing advanced support for internal AI teams in areas like model parallelism and tensor parallelism.
  • Assisting in and guiding the development of containerized pipeline orchestration infrastructure, including developing and distributing stable base container images, providing advanced profiling and observability, and updating internally maintained versions of deep learning frameworks and their companion libraries like CUDA, cuTile, cuDNN, NCCL, RDMA, TensorFlow, PyTorch, TorchRec, Flash Attention, PyTorch Lightning, and more.

Feature Engineering

  • Exploring and innovating within the online, offline, and nearline spaces at scale (millions of QPS, multi-terabytes of data, etc.).
  • Developing and refining the infrastructure necessary to transform raw data into valuable feature insights.
  • Utilizing leading open-source technologies like Spark, Beam, and Flink to process and structure feature data, ensuring its optimal storage in the Feature Store and serving feature data with high performance.

Model Serving Infrastructure

  • Building low-latency, high-performance applications serving very large and complex models across LLM and personalization models.
  • Building compute-efficient infrastructure on top of native cloud.
  • Enabling GPU-based inference for a large variety of use cases and CUDA-level optimizations for high performance.
  • Enabling on-device and online training.
  • Addressing challenges including scale (tens of thousands of QPS, multiple terabytes of data, billions of model parameters), agility (experimenting with hundreds of new ML models per quarter using thousands of features), and enabling GPU inference at scale.

ML Ops

  • Responsibility for the infrastructure that runs MLOps and experimentation systems across LinkedIn, including ramping, observability, and more.
  • Building tools that enable product and infrastructure engineers to optimize their models and deliver the best performance possible.

Qualifications

Basic Qualifications

  • Bachelor’s Degree in Computer Science or related technical discipline, or equivalent practical experience.
  • 4+ years of experience in the industry with leading/building deep learning systems.
  • 4+ years of experience with Java, C++, Python, Go, Rust, C#, and/or functional languages such as Scala or other relevant coding languages.
  • Hands-on experience developing distributed systems or other large-scale systems.

Preferred Qualifications

  • BS and 8+ years of relevant work experience, MS and 7+ years of relevant work experience, or PhD and 4+ years of relevant work experience.
  • Previous experience working with geographically distributed co-workers.
  • Outstanding interpersonal communication skills (including listening, speaking, and writing) and ability to work well in a diverse, team-focused environment with other SRE/SWE Engineers, Project Managers, etc.
  • Experience building ML applications, LLM serving, GPU serving.
  • Experience with search systems or similar large-scale distributed systems.
  • Expertise in machine learning infrastructure, including technologies like MLFlow, Kubeflow, and large-scale distributed systems.
  • Experience with distributed data processing engines like Flink, Beam, Spark, etc., feature engineering.
  • Co-author or maintainer of any open-source projects.
  • Familiarity with containers and container orchestration systems.
  • Expertise in deep learning frameworks and tensor libraries like PyTorch, TensorFlow, JAX/FLAX.

Skills

  • Distributed systems and Backend Systems Infrastructure
  • Java/Golang/Rust/Python
  • ML Algorithm Development, Machine Learning, and Deep Learning
  • Information Retrieval, Recommendation Systems, Distributed Serving, and Big Data
  • Communication and Stakeholder Management

Benefits

We strongly believe in the well-being of our employees and their families. That is why we offer generous health and wellness programs and time away for employees of all levels.

Pay

The pay range for this role is $175,000 - $287,000. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to skill set, depth of experience, certifications, and specific work location. This may be different in other locations due to differences in the cost of labor. The total compensation package for this position may also include annual performance bonus, stock, benefits, and/or other applicable incentive compensation plans.

Similar jobs