Technical Architect - MLE
Jobgether · United States · 1 wk ago
RemoteRemoteEngineeringFull-time
Accountabilities
- Architect and develop end-to-end agentic AI systems and multi-agent workflows from concept to production.
- Design agent architectures, including orchestration layers, communication protocols, agent roles, task planning mechanisms, and collaboration frameworks using technologies such as CrewAI, LangGraph, and AutoGen.
- Build advanced agent capabilities through custom tools, agent skills, integrations, and domain-specific workflows.
- Develop state management, memory systems, and context engineering approaches that enable persistent, reliable, and intelligent agent interactions.
- Deploy, scale, and maintain production-grade AI solutions across major cloud platforms including AWS, Google Cloud Platform, and Azure.
- Implement MLOps best practices, including monitoring, CI/CD, reliability improvements, and operational excellence.
- Integrate and optimize large language models using techniques such as Retrieval-Augmented Generation (RAG), prompt engineering, and parameter-efficient fine-tuning (PEFT).
- Create and maintain enterprise tool libraries, API integrations, database connections, and external service integrations for AI agents.
- Build observability and evaluation frameworks to measure agent performance, reliability, accuracy, cost, and latency.
- Implement monitoring and tracing solutions using tools such as LangSmith, Arize AI, or custom telemetry platforms.
- Establish quality metrics, evaluation processes, safety controls, and performance standards for production AI systems.
- Collaborate with technical teams and stakeholders to deliver innovative AI solutions while mentoring engineers and promoting technical excellence.
Requirements
- 8-12 years of professional experience in machine learning engineering, AI architecture, or related technical roles, with demonstrated experience delivering ML systems into production.
- Proven expertise designing and implementing multi-agent systems and agentic AI workflows.
- Strong programming skills in Python and experience with machine learning frameworks such as TensorFlow, PyTorch, and Transformers.
- Experience developing scalable applications using FastAPI, asynchronous programming, and microservices architectures.
- Hands-on experience building RAG systems and working with vector databases including Pinecone, Weaviate, or ChromaDB.
- Strong knowledge of LLM application monitoring, evaluation, and observability tools such as LangSmith, Weights & Biases, or similar platforms.
- Production-level experience with at least one major cloud platform: AWS, Google Cloud Platform, or Azure.
- Knowledge of cloud infrastructure including compute services, serverless functions, container orchestration platforms, and managed AI/ML services.
- Strong DevOps expertise including Infrastructure as Code tools such as Terraform or CloudFormation, CI/CD pipelines, Docker, and Kubernetes.
- Familiarity with distributed systems, message queues, event-driven architectures, and stateful AI orchestration patterns.
- Experience with AI evaluation methodologies, including trajectory analysis, tool-use validation, regression testing, and LLM-based evaluation frameworks.
- Understanding of AI safety practices, including data handling, guardrails, access controls, and secure prompt engineering.
- Knowledge of model lifecycle management, including routing strategies, model versioning, fallbacks, and optimization techniques.
- Strong problem-solving and analytical skills with the ability to solve complex technical challenges.
- Excellent communication skills with the ability to explain advanced AI concepts to both technical and non-technical audiences.
- Able to work independently, lead large-scale initiatives, and mentor other engineers.
- Experience with agile methodologies, software development lifecycle practices, and version control systems such as Git.
Benefits
- Remote work opportunity within the United States.
- Opportunity to work on cutting-edge AI, machine learning, and cloud transformation initiatives.
- Collaborative environment focused on innovation, learning, and professional growth.
- Exposure to advanced generative AI, agentic AI, and enterprise-scale technology solutions.
- Opportunities to collaborate with global teams and industry-leading technology partners.
- Career development opportunities within a rapidly growing AI-focused organization.
- Inclusive culture built around transparency, diversity, integrity, and continuous learning.
- Opportunity to contribute to impactful solutions addressing complex business challenges.