Jobs · Engineering · New York

Senior Site Reliability Engineer

Adobe · New York, NY · 2 days ago
Engineering$178k–$258k/yrFull-time

About the role

The Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and businesses access to hundreds of millions of curated, royalty-free assets that integrate directly with Adobe's desktop, mobile, and web apps, as well as Behance's creative community and Content's advanced services. Our infrastructure spans multiple cloud platforms and is orchestrated by cloud-native, containerized systems as we build the framework for Adobe's next wave of AI-powered products. The team blends startup energy with the resources of a large software company.

Responsibilities

  • Design, build, and operate large-scale, distributed, fault-tolerant systems, and the IaC, CI/CD, and automation tooling behind them
  • Own patch, vulnerability, and golden-image lifecycle management across the fleet — triage, remediate, automate
  • Contribute to multi-quarter initiatives: compute rationalization, cloud re-platforming, ML platform migration
  • Build and operate ML inference infrastructure — model serving, GPU workloads, language model gateway and routing — contributing to a broader systems portfolio
  • Help set infrastructure standards and security guardrails for agentic AI, and build the agent tooling we used to run our own operations
  • Partner across software, ML, and platform teams to bake reliability in from the start
  • Share on-call and fix challenges wherever problems show up — web services, databases, pipelines, ML systems

Requirements

  • Must have: BSc in Computer Science or equivalent experience in a practical setting
  • Python, plus familiarity with PHP, Node.js, or Ruby
  • Familiarity with deploying/operating ML inference pipelines (SageMaker, OpenAI, Bedrock, or equivalent) in production
  • Experience operating cloud-native compute over broad, scalable infrastructures — EC2 Auto Scaling Groups and Kubernetes in AWS; Azure/GCP a plus.
  • Experience with vulnerability/patch management and AMI/golden-image lifecycle automation at fleet scale
  • Strong debugging skills on distributed systems
  • Willingness to participate in on-call rotation

Nice to have

  • Production ML operations experience at any scale
  • GPU-backed compute and cost/performance tuning experience
  • Exposure to agentic AI workflows — LangGraph or similar, LLM gateway/routing across providers
  • Relational database operations experience (upgrades, IAM auth, performance tuning)

Technologies we use

  • Cloud & Orchestration: AWS (primary), Azure, GCP, Kubernetes, LambdaEdge & Security: Fastly, Datadome, WAF
  • ML/AI: SageMaker, Bedrock, vLLM, LangGraph, LangSmith, LLM gateway, MCP
  • Data & Caching: Aurora PostgreSQL, Memcached
  • IaC & Automation: Terraform, Terragrunt, Atlantis, Chef, Ansible, SSM, Docker
  • Languages: Python, PHP, bash, Ruby
  • CI/CD: Jenkins, Argo
  • Observability: New Relic, Splunk, Grafana, Prometheus

What you need to succeed

  • You're a self-starter — passionate, motivated, and comfortable working remotely
  • Some SDE background helps, but a willingness to build that skill matters more
  • You understand how software and systems come together and operate internet-hosted applications
  • You should also know your way around the infrastructure toolchain, including provisioning, config management, patch/image management, and container orchestration
  • You don't need to possess the qualifications of an ML engineer, but you should understand ML capabilities and how inference pipelines and model-serving systems work

Expected Pay Range

The expected pay range for this position is $139,000 -- $257,550 annually. Pay within this range varies by work location and may also depend on job-related knowledge, skills, and experience.

Similar jobs