Jobs · Engineering · New York

Lead Software Engineer: (Vice President), Cloud Infrastructure & Platform Engineering (Java, Agentic AI)

JPMorganChase · New York, NY · 4 wk ago
On-siteEngineeringFull-time

Job Responsibilities

  • Lead the design, build, and day-to-day operation of cloud platform capabilities and production environments at scale, emphasizing resiliency, security, and reliability
  • Build and maintain secure Java (Java 17+) and Spring Boot services, including RESTful APIs, with strong code review and debugging
  • Implement Infrastructure-as-Code / Everything-as-Code using Terraform, including secure provisioning and policy controls where applicable
  • Enable and support safe release practices (e.g., rolling and blue/green deployments) for high-availability, high-scale services
  • Improve operational stability by eliminating recurring issues and automating remediation/runbooks where appropriate
  • Apply SRE-aligned practices (SLIs/SLOs, error budgets, incident response) and participate in formal change and incident management processes
  • Use observability and APM tooling (e.g., Splunk, Datadog, Dynatrace) to monitor, troubleshoot, and improve production services
  • Deliver CI/CD in a “you build it, you own it” culture using Git and enterprise pipelines (e.g., Jules, Spinnaker, or similar)
  • Prototype and deliver AI-enabled/agentic workflows (e.g., tool-using agents, orchestration patterns, evaluation and guardrails) and translate stakeholder needs into technical designs
  • Mentor engineers and partner with application teams to drive best practices across reliability, performance, risk, cybersecurity, and legal/compliance requirements

Required Qualifications, Capabilities, And Skills

  • Formal training or certification on software engineering concepts and 5+ years of applied experience
  • Hands-on Java/J2EE development experience, with strong depth in Spring Boot and REST API design
  • AWS platform/infrastructure engineering experience, including designing and operating cloud workloads
  • Hands-on experience operating Kubernetes in high-availability environments, including deployments and incident troubleshooting
  • Strong Infrastructure-as-Code experience with Terraform (modules, state management, secure provisioning practices)
  • Production support experience using observability/APM tooling (e.g., Splunk, Datadog, Dynatrace), including root-cause analysis and reliability improvements
  • Experience with testing practices and frameworks (e.g., JUnit, Mockito) and performance testing tools (e.g., JMeter/BlazeMeter or equivalent)
  • Solid grounding in domain-driven design, microservices patterns, resiliency patterns, and operating services under change/incident management controls
  • Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security
  • Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices

Preferred Qualifications, Capabilities, And Skills

  • Experience with AWS services commonly used in modern platforms (e.g., ELB/ALB, Route 53, API Gateway, Lambda, ECS, managed caching such as Redis)
  • Experience implementing policy-as-code and guardrails for cloud/Kubernetes environments
  • Experience building platform “golden paths” (templates, paved roads, reusable modules) that accelerate application teams safely
  • Experience with secrets management tooling and patterns (e.g., rotation, least privilege, auditability) in regulated environments
  • Practical experience building or integrating AI/agentic workflows with measurable evaluation, safety controls, and monitoring in production-like environments
  • Operate Kubernetes platforms (preferably Amazon EKS), including cluster operations, upgrades, deployment strategies, networking/ingress, configuration and secrets management, and troubleshooting in high-availability environments

Similar jobs