Lead Software Engineer: (Vice President), Cloud Infrastructure & Platform Engineering (Java, Agentic AI)
JPMorganChase · New York, NY · 4 wk ago
On-siteEngineeringFull-time
Job Responsibilities
- Lead the design, build, and day-to-day operation of cloud platform capabilities and production environments at scale, emphasizing resiliency, security, and reliability
- Build and maintain secure Java (Java 17+) and Spring Boot services, including RESTful APIs, with strong code review and debugging
- Implement Infrastructure-as-Code / Everything-as-Code using Terraform, including secure provisioning and policy controls where applicable
- Enable and support safe release practices (e.g., rolling and blue/green deployments) for high-availability, high-scale services
- Improve operational stability by eliminating recurring issues and automating remediation/runbooks where appropriate
- Apply SRE-aligned practices (SLIs/SLOs, error budgets, incident response) and participate in formal change and incident management processes
- Use observability and APM tooling (e.g., Splunk, Datadog, Dynatrace) to monitor, troubleshoot, and improve production services
- Deliver CI/CD in a “you build it, you own it” culture using Git and enterprise pipelines (e.g., Jules, Spinnaker, or similar)
- Prototype and deliver AI-enabled/agentic workflows (e.g., tool-using agents, orchestration patterns, evaluation and guardrails) and translate stakeholder needs into technical designs
- Mentor engineers and partner with application teams to drive best practices across reliability, performance, risk, cybersecurity, and legal/compliance requirements
Required Qualifications, Capabilities, And Skills
- Formal training or certification on software engineering concepts and 5+ years of applied experience
- Hands-on Java/J2EE development experience, with strong depth in Spring Boot and REST API design
- AWS platform/infrastructure engineering experience, including designing and operating cloud workloads
- Hands-on experience operating Kubernetes in high-availability environments, including deployments and incident troubleshooting
- Strong Infrastructure-as-Code experience with Terraform (modules, state management, secure provisioning practices)
- Production support experience using observability/APM tooling (e.g., Splunk, Datadog, Dynatrace), including root-cause analysis and reliability improvements
- Experience with testing practices and frameworks (e.g., JUnit, Mockito) and performance testing tools (e.g., JMeter/BlazeMeter or equivalent)
- Solid grounding in domain-driven design, microservices patterns, resiliency patterns, and operating services under change/incident management controls
- Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security
- Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices
Preferred Qualifications, Capabilities, And Skills
- Experience with AWS services commonly used in modern platforms (e.g., ELB/ALB, Route 53, API Gateway, Lambda, ECS, managed caching such as Redis)
- Experience implementing policy-as-code and guardrails for cloud/Kubernetes environments
- Experience building platform “golden paths” (templates, paved roads, reusable modules) that accelerate application teams safely
- Experience with secrets management tooling and patterns (e.g., rotation, least privilege, auditability) in regulated environments
- Practical experience building or integrating AI/agentic workflows with measurable evaluation, safety controls, and monitoring in production-like environments
- Operate Kubernetes platforms (preferably Amazon EKS), including cluster operations, upgrades, deployment strategies, networking/ingress, configuration and secrets management, and troubleshooting in high-availability environments