Jobs · Engineering

LLM / GenAI Engineer

Evlo AI · Chicago, IL · 3 wk ago
RemoteRemoteEngineeringFull-time

About the role

The role is for someone who has moved beyond basic prompting and understands what it takes to build production-grade AI systems: robust RAG pipelines, agentic workflows, fine-tuning pipelines, and systematic evaluation frameworks. The engineering team owns complex pieces of a rapidly scaling AI platform, working directly with applied scientists and backend engineers to solve hard problems in generative AI.

Responsibilities

  • Design and implement production-grade RAG pipelines using LangChain, LlamaIndex, or custom architectures
  • Build and optimize vector database integrations such as Pinecone, Weaviate, or pgvector for semantic search at scale
  • Develop systematic LLM evaluation frameworks including benchmark suites, LLM-as-judge pipelines, and regression testing
  • Run instruction fine-tuning and parameter-efficient fine-tuning workflows like LoRA and QLoRA on domain-specific datasets
  • Deploy, monitor, and scale LLM applications using cloud infrastructure, containerization, and robust API design
  • Write clean, testable, well-documented code and actively participate in architecture reviews and team engineering standards

Requirements

  • 3-6 years of software engineering experience, with at least 2 years focused specifically on LLMs and generative AI in production
  • Deep familiarity with LLM orchestration frameworks such as LangChain or LlamaIndex
  • Solid understanding of embedding models, vector databases, and semantic similarity in production environments
  • Strong Python skills, with comfort in async programming, REST API design, and cloud infrastructure

Skills

  • Experience with model quantization, custom CUDA kernels, or contributing to open-source AI projects (bonus)

Similar jobs