Senior Lead Software Engineer - Agentic AI SRE Platform
hackajob · Seattle, WA · 3 wk ago
On-siteEngineeringFull-time
About the role
Join JPMorganChase’s AI/ML Data Platforms organization as a Senior Lead Software Engineer and pioneer the architecture of production-grade agentic AI and multi-agent systems for the Reliability Engineering team. You will design autonomous agents that integrate across telemetry pipelines, application stacks, and cloud infrastructure to drive engineering productivity, system reliability, and operational excellence at enterprise scale.
Responsibilities
- Design and develop enterprise AI systems that enhance engineering productivity, system reliability, and operational excellence.
- Architect autonomous agents that reason, analyze, and act across system and application reliability—not just integrating workflows.
- Architect and govern agentic system components including orchestration, memory/state management, workflow control, and human-in-the-loop patterns where needed.
- Own the agentic architecture’s non-functional requirements posture end-to-end: observability, security, guardrails, and evaluation.
- Provision and manage scalable AI infrastructure, deployment pipelines, and cost optimization.
- Execute creative software solutions, design, development, and technical troubleshooting with the ability to think beyond routine or conventional approaches.
- Develop secure, high-quality production code, and review and debug code written by others.
- Identify opportunities to eliminate or automate remediation of recurring issues to improve overall operational stability.
- Lead evaluation sessions with external vendors, startups, and internal teams to drive outcomes-oriented probing of architectural designs, technical credentials, and applicability for use within existing systems.
- Contribute to a team culture of diversity, opportunity, inclusion, and respect.
Requirements
- Formal training or certification on software engineering concepts and 5+ years of applied experience.
- Advanced proficiency in programming with Python.
- Proficiency in all aspects of the Software Development Life Cycle and in automation and continuous delivery methods.
- Hands-on experience delivering agentic AI to production (not just LLM wrappers or prompt-to-workflow integrations).
- Experience with AI agent frameworks (e.g., LangGraph, CrewAI), runtimes, orchestration, evaluation/benchmarking systems, and model selection policies.
- Experience developing production generative AI applications at scale with experimentation, model quality measurement, retrieval-augmented generation, and vector search.
- Strong cloud infrastructure experience (e.g., AWS) and infrastructure-as-code (e.g., Terraform).
- Experience building distributed systems or high-scale systems where reliability matters.
Preferred Qualifications
- Previous experience working within infrastructure, platform engineering, or reliability engineering domains.
- Knowledge of site reliability engineering concepts, including Service Level Objectives/Indicators, error budgets, incident management, resilience patterns, and observability.
- Previous experience as a machine learning software engineer in a dynamic technology company or startup.
Benefits
- Comprehensive health care coverage.
- On-site health and wellness centers.
- Retirement savings plan.
- Backup childcare.
- Tuition reimbursement.
- Mental health support and financial coaching.
- Discretionary incentive compensation, paid in cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions.
Pay
Competitive total rewards package including base salary determined based on the role, experience, skill set, and location. Additional details about total compensation and benefits will be provided during the hiring process.