Technical Architect - ML - GenAI
Quantiphi · United States · 3 wk ago
RemoteRemoteEngineeringFull-time
About the role
We are looking for a Generative AI Architect / Lead to design and deliver enterprise-grade GenAI solutions using AWS Bedrock and Agentcore. This role focuses on building scalable applications leveraging large language models (LLMs), retrieval-augmented generation (RAG), and agentic AI workflows. The ideal candidate will be a hands-on architect who can define solution architecture, guide teams, and actively contribute to development while ensuring performance, scalability, and cost efficiency.
Responsibilities
- Design and implement GenAI solutions using AWS Bedrock and Agentcore
- Define architecture for LLM-based applications, including RAG pipelines and agentic workflows
- Develop and orchestrate agentic AI workflows, enabling multi-step reasoning, tool usage, and task automation
- Build and manage RAG pipelines, including embeddings, retrieval mechanisms, and vector databases
- Integrate LLM capabilities into enterprise applications via APIs and backend services
- Design and optimize prompt engineering strategies for accuracy, relevance, and performance
- Work with structured and unstructured data sources to enable knowledge-driven AI applications
- Ensure model evaluation, monitoring, and optimization for latency, cost, and response quality
- Collaborate with application, data, and platform teams for end-to-end solution delivery
- Define best practices for security, governance, and responsible AI usage
- Troubleshoot and resolve issues in production GenAI systems
- Provide technical leadership and mentor team members while remaining hands-on
- Develop and maintain Model Context Protocol (MCP) implementations to manage state, context windows, memory, and prompt orchestration across distributed agent systems
Requirements
- 8+ years of relevant hands-on technical experience implementing and developing cloud ML solutions on AWS
- Hands-on experience with AWS services, including AWS Sagemaker and Bedrock, leveraging different types of data sources, training jobs, real-time and batch applications
- Design and implement agentic AI architectures using frameworks such as LangChain, Strand Agents, etc., enabling autonomous task planning, decision-making, and multi-step reasoning
- Hands-on experience with Amazon AgentCore for building, deploying, and scaling production-grade agentic AI applications, including agent memory management, tool registry, and observability
- Architect and deploy scalable AI solutions on AWS, leveraging services like Lambda, Bedrock, Step Functions, S3, API Gateway, and SageMaker
- Proficiency in working with LLM APIs (e.g., Claude, Nova, and other third-party LLM providers), including API integration and multi-model orchestration strategies
- Hands-on experience fine-tuning or optimizing large language models (LLM)
- Familiarity with LLM tool use, prompt templating, and context management
- Strong expertise in Vector Databases, including indexing strategies, embedding generation, similarity search, and integration with RAG architectures
- Model Evaluation & Optimization: Evaluate LLM's zero-shot and few-shot capabilities, fine-tuning hyperparameters, ensuring task generalization, and exploring model interpretability for robust web app integration
- Experience with at least one of the workflow orchestration tools: Airflow, StepFunctions, SageMaker Pipelines, Kubeflow, etc.
- Experience implementing secure, scalable APIs and integrating with 3rd-party data sources and tools
- Ability to collaborate with cross-functional teams such as Developers, QA, Project Managers, and other stakeholders to understand their requirements and implement solutions
- Experience with Deep Learning Concepts: Transformers, BERT, Attention models, tokenization, embeddings
Nice to Have
- Experience with software development, exposure to frontend/backend frameworks, and communication protocols
- Experience working on Infrastructure as Code (IaC) and CI/CD pipelines
- Experience with NLP concepts: syntactic/semantic analysis, NER, etc.
Work location: Remote (US)