Jobs · Information Technology · California

GENAI ENGINEER – LLM INFRASTRUCTURE & INFERENCE SERVICES

VeriiPro · Santa Clara, CA · Today
Information TechnologyContract
We are looking for a GenAI Engineer with strong expertise in LLM infrastructure, model deployment, and high-performance inference services. The ideal candidate will build and manage scalable enterprise GenAI platforms across GPU infrastructure and cloud environments. Key Responsibilities Deploy, host, and manage Large Language Models (LLMs) on GPU infrastructure for production environments.Build scalable, high-performance inference services using vLLM, TensorRT-LLM, Triton Inference Server, and Ray Serve.Optimize model serving for latency, throughput, GPU utilization, and cost efficiency.Develop AI platform services and APIs using Python, FastAPI, Microservices, and Kubernetes.Implement RAG pipelines, vector databases, and agentic AI frameworks such as LangChain and LangGraph.Manage GPU infrastructure, containerization, and cloud deployments across AWS, Azure, or GCP.Establish MLOps/LLMOps practices including CI/CD, model deployment, monitoring, observability, and governance.Perform performance tuning, benchmarking, capacity planning, and production support for enterprise GenAI platforms.Collaborate with architects, data scientists, and product teams to deliver scalable, secure, and reliable AI solutions. Core Technologies LLM: vLLM, TensorRT-LLM, Triton Inference Server, Ray ServeAI/GenAI: RAG, LangChain, LangGraph, Vector DatabasesDevelopment: Python, FastAPI, MicroservicesInfrastructure: Kubernetes, Docker, GPU InfrastructureCloud: AWS, Azure, GCPMLOps/LLMOps: CI/CD, Monitoring, Observability, Model Deployment, Governance

Similar jobs