Jobs · Marketing · California

Senior Production Operations Engineer

Sony Interactive Entertainment · San Diego, CA · 1 wk ago
HybridMarketing$172k–$258k/yrFull-time

Sony Interactive Entertainment isn’t just the Best Place to Play — it’s also the Best Place to Work. As the company behind the PlayStation brand and a subsidiary of Sony Group Corporation, SIE is a dynamic technology leader delivering cutting-edge hardware and network services to over 100 million users. We’re also an entertainment powerhouse, home to some of the world’s most beloved and recognizable intellectual properties. Our mission is to create and nurture the experiences under the PlayStation brand, synonymous with entertainment excellence and creativity.

As a member of the Production Operations Engineering team within the platform technology group, you will ensure key user experiences on the platform remain available, resilient, and high-performing across time zones and critical business periods. You’ll enable service teams to deliver new products and technical features while responding to production incidents, including rotational on-call support and urgent operational escalations. This role empowers you to drive technical initiatives, identify improvements in process and technology, and enhance production operations supporting millions of users.

Responsibilities

  • Own application operations and production support for internal and public-facing services in AWS and container environments, focusing on availability, resiliency, scalability, performance, and security.
  • Drive production readiness for new services and features, including provisioning, automation, monitoring, alerting, dashboards, runbooks, rollback planning, and operational acceptance criteria.
  • Build scripts, tools, and repeatable workflows to reduce toil, improve incident response, and standardize operational practices across the environment.
  • Improve observability and event correlation using actionable service health signals, dashboards, alerts, and incident workflows to reduce mean time to detect (MTTD) and mean time to resolve (MTTR).
  • Support and enhance incident management practices using tools such as Slack, JIRA, ServiceNow, BigPanda, and related escalation/notification systems.
  • Partner with SRE, data services, CI/CD, service engineering, platform hosting, and product teams to improve end-to-end reliability and operational readiness.
  • Drive performance, capacity, and cost optimization for services using AWS, Kubernetes, EKS/UKS, autoscaling, and related cloud patterns.
  • Leverage AI-assisted and agentic engineering tools (e.g., Codex, Claude, Cursor) to accelerate development, operational automation, and production-signal analysis.
  • Provide rotational on-call support, lead incident triage/mitigation, and document post-incident reviews with corrective actions and prevention opportunities.

Qualifications

  • Senior hands-on engineer equally comfortable with software development, systems engineering, and production operations.
  • Proven ability to operate distributed services at scale while automating toil, improving processes, and raising production readiness standards.
  • Strong troubleshooting skills across user experience, application runtime, Linux, infrastructure, cloud, Kubernetes/EKS/UKS, and network layers.
  • Highly independent technical leader who uses sound judgment, mentors less experienced engineers, and drives follow-through during incidents.

Skills

  • Distributed service operations, Unix/Linux systems internals, networking, TCP/IP, HTTP/HTTPS, DNS, load balancing, and API troubleshooting.
  • AWS service operations, including ALB, Route 53, API Gateway, Lambda, RDS, DynamoDB, ElastiCache, and Java/API services.
  • Containers and orchestration, including Docker, Kubernetes, EKS/UKS, and Fargate.
  • Automation and software development in Python, Go, or Java; source control and configuration management using GitHub/Git, Ansible, Chef, or similar.
  • Infrastructure as Code and CI/CD using Terraform, CloudFormation, Jenkins, Spinnaker, or comparable tooling.
  • Observability, incident management, and collaboration tools such as Datadog, CloudWatch, Splunk, Grafana, BigPanda, Slack, JIRA, and ServiceNow.
  • Strong communication skills, including the ability to adjust messaging for different audiences, translate complex technical issues, and articulate business impact.
  • Data reporting, analytics, or operational data platforms (e.g., SQL, MySQL, Oracle, Snowflake, or big data systems).
  • Experience with AI-assisted development or agentic engineering workflows using tools like Codex, Claude, or Cursor.

Experience

  • BS degree or equivalent in Computer Science, Software Engineering, or a related technical field.
  • 5+ years operating and supporting services in production environments at scale.
  • 3+ years of AWS Cloud experience deploying, tuning, and operating Java/API services.
  • Hands-on incident management experience, including on-call, incident coordination, stakeholder communications, RCA, and corrective-action tracking.
  • Experience with payment processing, commerce, entitlement, wallet, fraud, or financial systems is a plus.

Pay

The estimated base pay range for this role is $172,100—$258,100 USD.

Benefits

  • Top-tier medical, dental, and vision coverage.
  • Matching 401(k) retirement plan.
  • Paid time off and wellness program.
  • Employee discounts on Sony products.
  • Eligibility for a bonus package.

Similar jobs

Senior Production Engineer

Mitsubishi Chemical AmericaLa Porte, TX· 1 mo ago
Management$130k–$140k/yrapply on mitsubishichemicalgroup.wd3.myworkdayjobs.com