Senior AI Platform Engineer – Agentic AI
Position Summary
We are seeking a highly skilled Senior AI Engineer to help design and build an enterprise-scale Agentic AI platform that enables multiple business domains to develop, deploy, monitor, and govern autonomous AI agents. This role goes beyond traditional LLM application development and requires hands-on expertise in agent orchestration, AI platform architecture, model governance, memory management, observability, cost attribution, multi-agent systems, and scalable cloud-native AI solutions. The ideal candidate will have experience building production-grade AI systems using Azure AI Foundry, LangChain, LangGraph, vector databases, API gateways, and modern AI engineering practices. The individual should be comfortable making architecture decisions, evaluating technology trade-offs, and designing enterprise-ready solutions that support security, scalability, monitoring, and cost control.
Key Responsibilities
- Agentic AI Solution Development
- Build sophisticated multi-agent AI systems for enterprise use cases.
- Implement supervisor-worker, sequential, orchestration, choreography, ReAct, Planner-Executor, and Writer-Critic agent architectures.
- Develop scalable agent communication and execution frameworks.
- Design closed-loop AI workflows with validation, retry, evaluation, and feedback mechanisms.
- Enterprise AI Platform Engineering
- Build reusable AI platform capabilities consumed by multiple business teams.
- Implement enterprise-grade AI governance and operational controls.
- Design API-driven AI Service Architecture With: Rate limiting, Quota management, Multi-tenant usage tracking, Cost attribution, Authentication & authorization, Audit logging.
- Enable structured onboarding and lifecycle management of AI agents.
- Multi-Agent Orchestration
- Design Orchestration Frameworks Where Agents Communicate Through: Direct calls, Event-driven architectures, Message queues, Publish-subscribe patterns.
- Evaluate technologies such as Kafka, Azure Durable Functions, Service Bus, and event-driven workflows.
- AI Memory & Knowledge Systems
- Design short-term and long-term memory architectures.
- Implement: Vector databases, Semantic caching, Conversation memory, Agent state persistence, Retrieval-Augmented Generation (RAG).
- Develop knowledge orchestration frameworks supporting agent collaboration.
- Ontology & Graph-based Intelligence
- Work with graph databases and enterprise knowledge models.
- Support ontology-driven AI applications.
- Build knowledge graphs that enable relationship-based reasoning and signal generation.
- Design systems that combine structured, unstructured, and graph-based knowledge sources.
- Model Governance & FinOps
- Implement AI consumption governance across business domains.
- Track: Token usage, Model consumption, API utilization, Operational costs.
- Create chargeback/showback mechanisms for enterprise teams.
- Support AI FinOps reporting and capacity planning.
- Reliability, Monitoring & Observability
- Design observability frameworks for AI applications.
- Monitor: Agent executions, Tool usage, Latency, Hallucinations, Failure rates, Model quality.
- Create dashboards and operational metrics for enterprise AI workloads.
- Responsible AI & Security
- Implement: Guardrails, Safety controls, Prompt protection, Data masking, PII protection, Human-in-the-loop validation.
- Ensure compliance with enterprise security and governance policies.
- Build secure agentic systems handling sensitive business data.
- AI Evaluation & Optimization
- Develop Frameworks For: Agent evaluation, Tool evaluation, Response quality measurement, Closed-loop evaluation, Hallucination detection.
- Apply Advanced AI Engineering Techniques Including: Context engineering, Prompt engineering, Retrieval optimization, Agent tuning, AI system benchmarking.
Qualifications
- 7+ years in software engineering or platform engineering.
- 3+ years building AI/ML or Generative AI solutions.
- Experience delivering enterprise-scale production AI applications.
- Experience designing AI architectures rather than only building individual AI applications.
- Technical Skills
- Generative AI & Agentic Frameworks
- Azure AI Foundry
- Azure OpenAI
- LangChain
- LangGraph
- Semantic Kernel (preferred)
- MCP (Model Context Protocol)
- Cloud Platforms
- Microsoft Azure (required)
- Experience with GCP or AWS is a plus
- Enterprise Integration
- API gateways and AI governance platforms
- Azure API Management (APIM)
- REST APIs
- Event-driven systems
- Programming
- Python (required)
- C# (.NET) preferred
- SQL
- Data & Storage
- Cosmos DB
- PostgreSQL
- MongoDB
- Vector databases
- Graph databases (Neo4j, Stardog, Neptune, etc.)
- Messaging & Streaming
- Kafka
- Azure Service Bus
- Event Grid
- Durable Functions
- AI Operations
- AI observability
- Monitoring & logging
- Token usage analysis
- Cost optimization
- Model lifecycle management
Preferred Qualifications
- Experience implementing ontology-driven solutions.
- Experience with enterprise knowledge graphs.
- Experience building autonomous AI systems.
- Experience with AI governance and responsible AI frameworks.
- Experience designing reusable AI platforms used by multiple business units.
- Experience with healthcare, financial services, insurance, or regulated industries.