Jobs · Engineering

LLM / GenAI Engineer

Evlo AI · Raleigh, NC · 3 wk ago
RemoteRemoteEngineeringFull-time

About the role

The role is for someone who has moved beyond basic prompting and understands what it takes to build production-grade AI systems: robust RAG pipelines, agentic workflows, fine-tuning pipelines, and systematic evaluation frameworks. The engineering team owns complex pieces of a high-scale AI platform, working directly with applied scientists and backend engineers to deploy reliable generative models.

Responsibilities

  • Design and implement production-grade RAG pipelines using LangChain, LlamaIndex, or custom orchestration architectures
  • Build and optimize vector database integrations including Pinecone, Weaviate, or pgvector for semantic search at scale
  • Develop systematic LLM evaluation frameworks incorporating benchmark suites, LLM-as-judge pipelines, and regression testing
  • Execute instruction fine-tuning and parameter-efficient fine-tuning techniques such as LoRA and QLoRA on domain-specific datasets
  • Integrate monitoring and observability tools like LangSmith or Arize to track token usage, latency, and output drift
  • Write clean, testable, and well-documented Python code while participating in rigorous architecture reviews

Requirements

  • 3–6 years of software engineering experience, with at least 2 years focused specifically on LLMs and generative AI in production
  • Deep familiarity with LLM orchestration frameworks, prompt engineering best practices, and embedding models
  • Strong Python skills with comfort in async programming, REST API design, and cloud infrastructure
  • Solid understanding of ML fundamentals, tokenization, context windows, and vector similarity search
  • Bachelor's or Master's degree in Computer Science, Artificial Intelligence, or equivalent practical experience

Skills

  • Bonus: Experience contributing to open-source AI projects or published research in NLP and generative modeling

Similar jobs