Jobs · Engineering · Tennessee

Senior Manager, AI Agent Software Engineering

Oracle · Nashville, TN · 6 days ago
Engineering$146k–$306k/yrFull-time

About the role

The Senior Manager will lead a team of engineers while serving as a hands-on technical leader responsible for defining, building, and operating next-generation AI systems on Oracle Cloud Infrastructure (OCI). This leader will set the architecture and engineering direction for production-grade agentic AI platforms, autonomous workflows, scalable inference infrastructure, and enterprise AI applications used in large-scale, business-critical environments. The role requires an experienced engineering manager who can build and develop high-performing teams, translate ambiguous product and platform goals into a durable technical strategy, and drive execution across multiple organizations.

The successful candidate will be accountable for hiring, mentoring, performance management, technical planning, and delivery, while remaining actively involved in system design, prototyping, coding, code reviews, operational readiness, and incident follow-up. The ideal candidate combines deep distributed-systems expertise with practical, hands-on experience building AI agents and AI-native applications, including developing and orchestrating LLM-based agents, tools, APIs, memory systems, retrieval pipelines, evaluations, guardrails, and cloud-service integrations.

Responsibilities

  • Lead and develop a team responsible for OCI AI platform capabilities, including agent execution, inference, orchestration, evaluation, and observability.
  • Set the technical direction for production-grade agentic AI systems that support reasoning, planning, tool use, multi-step workflows, and human escalation.
  • Remain hands-on in architecture, prototyping, coding, debugging, and code reviews for critical AI-agent components.
  • Guide the development of services for tool calling, memory, context management, MCP integration, retrieval, multi-agent coordination, policy enforcement, and evaluation.
  • Own delivery across distributed systems optimized for reliability, performance, security, cost, and multi-tenant operation.
  • Translate broad goals into roadmaps, staffing plans, milestones, and measurable outcomes.
  • Partner across infrastructure, security, data, product, and application teams to drive execution.
  • Establish AgentOps and LLMOps practices for tracing, monitoring, testing, safety guardrails, versioning, and production readiness.
  • Recruit, coach, and retain engineers while managing performance and developing senior technical leaders.
  • Own production outcomes, including reliability, security, cost efficiency, supportability, and delivery predictability.

Requirements

  • Bachelor's, Master's, or Ph.D. in Computer Science, AI/ML, Engineering, or a related field, or equivalent experience.
  • 8+ years of software engineering experience, including ownership of production systems.
  • 2+ years of engineering management experience, including hiring, coaching, performance management, and delivery ownership.
  • Proven ability to lead teams while remaining technically engaged in design, coding, reviews, debugging, and operations.
  • Deep experience with distributed systems, cloud platforms, or AI/ML infrastructure.
  • Hands-on experience building AI agents, autonomous workflows, tool-using systems, or multi-step orchestration.
  • Experience with frameworks such as LangGraph, LangChain, CrewAI, AutoGen, LlamaIndex, or similar tools.
  • Strong understanding of LLM patterns, including tool calling, RAG, memory, context management, evaluation, and safety.
  • Strong Python skills and experience with Kubernetes, Docker, observability, scalability, and fault tolerance.
  • Strong understanding of AI security, governance, access control, auditability, and operational risk.
  • Excellent communication and cross-functional leadership skills.

Preferred Qualifications

  • Experience managing teams that build AI platforms, agent runtimes, inference systems, or developer platforms.
  • Experience with GPU inference optimization, model serving, workflow engines, or multi-tenant cloud services.
  • Experience integrating AI systems with enterprise APIs, databases, identity systems, vector stores, and policy layers.
  • Experience with agent evaluation, adversarial testing, regression gates, and production observability.
  • Experience using AI-assisted development tools such as Codex, Claude Code, Cursor, or Copilot.
  • Experience in enterprise, cloud infrastructure, regulated, or mission-critical environments.
  • Experience in data science and applied machine learning, including classical ML techniques, deep learning models, model evaluation, and production deployment.

Pay

Hiring range in USD from: $146,300 - $306,400 per year. May be eligible for bonus, equity, and compensation deferral.

Benefits

  • Medical, dental, and vision insurance, including expert medical opinion.
  • Short-term and long-term disability insurance.
  • Life insurance and AD&D.
  • Supplemental life insurance (Employee/Spouse/Child).
  • Health care and dependent care Flexible Spending Accounts.
  • Pre-tax commuter and parking benefits.
  • 401(k) Savings and Investment Plan with company match.
  • Paid time off: Flexible Vacation for salaried employees; accrued vacation for others (13 days annually for the first three years, 18 days annually thereafter, prorated for part-time).
  • 11 paid holidays.
  • Paid sick leave: 72 hours upon hire, refreshes annually, carries over up to 112 hours.
  • Paid parental leave.
  • Adoption assistance.
  • Employee Stock Purchase Plan.
  • Financial planning and group legal services.
  • Voluntary benefits including auto, homeowner, and pet insurance.

Similar jobs