Staff Production Operations Engineer
About the role
As a member of the Production Operations Engineering team within the platform technology group at Sony Interactive Entertainment (SIE), you will help keep key user experiences on the PlayStation platform available, resilient, and high-performing across time zones and critical business periods. You will enable service teams to deliver new products and technical features while responding to production service needs, including rotational on-call, incident response, and urgent operational escalations. This Staff individual contributor role carries broad technical ownership, empowering you to define and lead technical initiatives that improve process, technology, and production operations for millions of users.
Responsibilities
- Lead application operations and production support for internal and public-facing services in AWS and container environments, focusing on availability, resiliency, scalability, performance, and security.
- Define and evolve production operations architecture, readiness standards, and reusable patterns for services/features, including automation, monitoring, runbooks, rollback plans, and operational acceptance criteria.
- Architect and scale tools, scripts, and automation frameworks across API/service telemetry, release validation, and operational readiness to reduce toil and improve incident response.
- Establish CI/CD operational gates, production validation practices, and governance mechanisms to improve release confidence, scalability, and engineering velocity.
- Advance observability, event correlation, and operational metrics using service health signals, automation health, release risk indicators, dashboards, and alerts to reduce mean time to detect (MTTD) and mean time to resolve (MTTR).
- Partner with SRE, data services, CI/CD, service engineering, platform hosting, product, and business teams to shift reliability left and improve shared ownership of production quality.
- Lead performance, capacity, infrastructure modernization, and cost optimization for services using AWS, Kubernetes, EKS/UKS, autoscaling, and related cloud patterns.
- Use AI-assisted and agentic tools (e.g., Codex, Claude, Cursor) to accelerate development, operational automation, and production-signal analysis.
- Provide rotational on-call support and improve incident workflows using Slack, JIRA, ServiceNow, BigPanda, and related escalation systems; lead root cause analysis (RCA) and systemic fixes for recurring production issues.
Qualifications
- Staff-level hands-on engineer equally comfortable with software development, systems engineering, production operations, and service architecture.
- System Architect and full-stack engineer who connects customer experience, application behavior, APIs, data flows, infrastructure, and operational controls into resilient designs.
- Proven ability to define operational standards, quality gates, reusable frameworks, and governance practices across teams.
- Independent technical leader who influences cross-functional partners, mentors engineers, and drives adoption of scalable practices.
- Strong systems-thinking and troubleshooting depth across runtime, cloud, Kubernetes/EKS/UKS, network, observability, release, and incident layers.
- BS degree or equivalent in Computer Science, Software Engineering, or a related technical area.
- 10+ years operating and supporting services in production environments at scale.
- 5+ years of AWS Cloud experience deploying, tuning, and operating Java/API services.
- Hands-on incident management experience, including on-call, incident coordination, stakeholder communications, RCA, and corrective-action tracking.
- Experience with payment processing, commerce, entitlement, wallet, fraud, or financial systems is a plus.
Skills
- Distributed service architecture and operations, Unix/Linux systems internals, networking, TCP/IP, HTTP/HTTPS, DNS, load balancing, and API troubleshooting.
- AWS service operations, including ALB, Route 53, API Gateway, Lambda, RDS, DynamoDB, ElastiCache, and Java/API services.
- Containers and orchestration, including Docker, Kubernetes, EKS/UKS, and Fargate.
- Automation and software development in Python, Go, or Java; source control and configuration management using GitHub/Git, Ansible, Chef, or similar.
- Infrastructure as Code and CI/CD using Terraform, CloudFormation, Jenkins, Spinnaker, or comparable tooling.
- Observability, incident management, and collaboration tooling such as Datadog, CloudWatch, Splunk, Grafana, BigPanda, Slack, JIRA, and ServiceNow.
- Strong communication skills, including the ability to adjust messaging by audience, translate complex technical issues for cross-functional partners, and articulate technical detail and business impact.
- Data reporting, analytics, or operational data platforms such as SQL, MySQL, Oracle, Snowflake, or big data systems.
- AI-assisted development or agentic engineering workflow experience using tools such as Codex, Claude, Cursor, or similar.
Pay
The estimated base pay range for this role is $199,400—$299,200 USD.
Benefits
- Top-tier medical, dental, and vision coverage.
- Matching 401(k) retirement plan.
- Paid time off and wellness program.
- Employee discounts on Sony products.
- Eligibility for a bonus package.