Site Reliability Engineer III- Production Management
Job Responsibilities
Guides and assists others in the areas of building appropriate level designs and gaining consensus from peers where appropriate, supporting adoption of site reliability engineering best practices within your team.
Collaborates with other software engineers and teams to design, develop, test, and implement deployment and reliability approaches using automated continuous integration and continuous delivery pipelines.
Uses enterprise-authorized AI capabilities within the work environment to accelerate incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
Continuously monitors operational signals (alerts, dashboards, service health indicators, and desk/user feedback) and drives rapid engagement, ownership, and coordination when issues emerge.
Enhance production observability and reliability by improving instrumentation, dashboards, and alert quality; use trend and incident data to reduce noise, prevent recurrence, and measurably lower MTTR and incident frequency.
Owns and improves the support operating model: maintains/runs runbooks, ensures clean handoffs and escalation paths, manages shift coverage, and drives disciplined post-incident follow-through (actions, owners, due dates).
Partners with engineering and adjacent service owners to drive root-cause fixes and stronger change/release standards, identifies cross-service dependencies early, and builds lightweight automation/self-service tooling to reduce manual steps and speed repeatable triage.
Leads L1/L2 production support using SRE practices: quickly triages incidents, investigates likely causes, applies safe mitigations/workarounds, and provides timely, clear stakeholder updates through to resolution.
Applies enterprise-authorized AI capabilities within the work environment to identify patterns in operational signals that indicate reliability risk or recurring toil, prioritizing reuse-first improvements tied to SLO outcomes.
Required Qualifications, Capabilities, And Skills
- Formal training or certification on site reliability engineering concepts and 3+ years applied experience
- Proficient in site reliability culture and principles and familiarity with how to implement site reliability within an application or platform
- Proficient in at least one programming language such as Python, Java/Spring Boot, and .Net
- Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows with strong validation habits and awareness of data sensitivity
- Able to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements
- Practical cloud experience in AWS supporting production services (visibility/troubleshooting, deployments, and operational hygiene)
- Experience operating in time-sensitive environments (market hours or similar), with strong incident discipline and stakeholder communication
- Strong debugging fundamentals across distributed systems (logs/metrics, latency analysis, dependency failures, data issues)
Preferred Qualifications, Capabilities, And Skills
- Prior Markets experience, especially pricing / risk / market data support (Fixed Income a plus)
- Familiarity with messaging/streaming patterns (Kafka/MQ) and data quality checks in execution and pricing pipelines
- Experience applying AI-assisted tooling to reduce support toil
About Us
JPMorganChase, one of the oldest financial institutions, offers innovative financial solutions to millions of consumers, small businesses and many of the world's most prominent corporate, institutional and government clients under the J.P. Morgan and Chase brands. Our history spans over 200 years and today we are a leader in investment banking, consumer and small business banking, commercial banking, financial transaction processing and asset management. We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process.
About The Team
J.P. Morgan's Commercial & Investment Bank is a global leader across banking, markets, securities services and payments. Corporations, governments and institutions throughout the world entrust us with their business in more than 100 countries. The Commercial & Investment Bank provides strategic advice, raises capital, manages risk and extends liquidity in markets around the world.