LLM / GenAI Engineer
Evlo AI · Washington, DC · Yesterday
RemoteRemoteEngineeringFull-time
About The Role The role is for someone who has moved beyond basic prompting and understands what it takes to build production-grade AI systems: robust RAG pipelines, agentic workflows, fine-tuning pipelines, and systematic evaluation frameworks. The engineering team owns complex pieces of a scalable AI platform, collaborating closely with applied scientists, backend engineers, and stakeholders to deliver high-performance generative AI solutions. Key Responsibilities Design and implement production-ready RAG architectures using LangChain, LlamaIndex, or custom frameworksBuild and optimize vector database integrations including Pinecone, Weaviate, and pgvector for low-latency semantic search at scaleDevelop systematic LLM evaluation frameworks incorporating benchmark suites, LLM-as-judge pipelines, and regression testingRun parameter-efficient fine-tuning pipelines such as LoRA and QLoRA on domain-specific datasets using PyTorch and Hugging FaceIntegrate Large Language Models with external tools and APIs to support complex, multi-step agentic workflowsWrite observable, tested, and well-documented Python code, participating actively in architecture reviews and team deployments What We Are Looking For 3–6 years of software engineering experience, with a minimum of 2 years specifically focused on building and deploying LLM and GenAI applications in productionDeep familiarity with LLM orchestration frameworks like LangChain or LlamaIndex and vector database managementStrong Python development skills including asynchronous programming, REST API design, and cloud infrastructure integrationSolid understanding of embedding models, tokenization, prompt engineering limits, and foundational transformer architecturesBachelor's degree in Computer Science, Artificial Intelligence, or equivalent practical experienceBonus: Experience with model quantization, vLLM optimization, TensorRT-LLM, or publishing research in NLP/GenAI domains