LLM / GenAI Engineer
Evlo AI · Chicago, IL · 3 wk ago
RemoteRemoteEngineeringFull-time
About the role
The role is for someone who has moved beyond basic prompting and understands what it takes to build production-grade AI systems: robust RAG pipelines, agentic workflows, fine-tuning pipelines, and systematic evaluation frameworks. The engineering team owns complex pieces of a rapidly scaling AI platform, working directly with applied scientists and backend engineers to solve hard problems in generative AI.
Responsibilities
- Design and implement production-grade RAG pipelines using LangChain, LlamaIndex, or custom architectures
- Build and optimize vector database integrations such as Pinecone, Weaviate, or pgvector for semantic search at scale
- Develop systematic LLM evaluation frameworks including benchmark suites, LLM-as-judge pipelines, and regression testing
- Run instruction fine-tuning and parameter-efficient fine-tuning workflows like LoRA and QLoRA on domain-specific datasets
- Deploy, monitor, and scale LLM applications using cloud infrastructure, containerization, and robust API design
- Write clean, testable, well-documented code and actively participate in architecture reviews and team engineering standards
Requirements
- 3-6 years of software engineering experience, with at least 2 years focused specifically on LLMs and generative AI in production
- Deep familiarity with LLM orchestration frameworks such as LangChain or LlamaIndex
- Solid understanding of embedding models, vector databases, and semantic similarity in production environments
- Strong Python skills, with comfort in async programming, REST API design, and cloud infrastructure
Skills
- Experience with model quantization, custom CUDA kernels, or contributing to open-source AI projects (bonus)