LLM / GenAI Engineer
Evlo AI · Raleigh, NC · 3 wk ago
RemoteRemoteEngineeringFull-time
About the role
The role is for someone who has moved beyond basic prompting and understands what it takes to build production-grade AI systems: robust RAG pipelines, agentic workflows, fine-tuning pipelines, and systematic evaluation frameworks. The engineering team owns complex pieces of a high-scale AI platform, working directly with applied scientists and backend engineers to deploy reliable generative models.
Responsibilities
- Design and implement production-grade RAG pipelines using LangChain, LlamaIndex, or custom orchestration architectures
- Build and optimize vector database integrations including Pinecone, Weaviate, or pgvector for semantic search at scale
- Develop systematic LLM evaluation frameworks incorporating benchmark suites, LLM-as-judge pipelines, and regression testing
- Execute instruction fine-tuning and parameter-efficient fine-tuning techniques such as LoRA and QLoRA on domain-specific datasets
- Integrate monitoring and observability tools like LangSmith or Arize to track token usage, latency, and output drift
- Write clean, testable, and well-documented Python code while participating in rigorous architecture reviews
Requirements
- 3–6 years of software engineering experience, with at least 2 years focused specifically on LLMs and generative AI in production
- Deep familiarity with LLM orchestration frameworks, prompt engineering best practices, and embedding models
- Strong Python skills with comfort in async programming, REST API design, and cloud infrastructure
- Solid understanding of ML fundamentals, tokenization, context windows, and vector similarity search
- Bachelor's or Master's degree in Computer Science, Artificial Intelligence, or equivalent practical experience
Skills
- Bonus: Experience contributing to open-source AI projects or published research in NLP and generative modeling