LLM Engineer
TALENT Software Services · Cincinnati, OH · Yesterday
RemoteRemoteOTHRFull-time
Job Details Job Title: LLM EngineerLocation: Cincinnati, OHWork Location: Remote - USADuration: 1 yearExperience Required: 8+ yearsRole Category: AI and AutomationEducation: Any degreeStart Date: 15-September-2026 Role Overview Design, build, optimize, deploy, and operate Large Language Model (LLM) and Small Language Model (SLM) capabilities. Build secure, reliable, reusable, and enterprise-ready AI capabilities. Support:Agentic AI workflowsAI for SDLCKnowledge retrievalModel evaluationPrivate AI hostingAgentOpsWork closely with:Principal AI ArchitectAI Engineering LeadPlatform EngineersSecurity teamsEnterprise ArchitectureProduct OwnersDomain teams Must-Have Technical Skills LLM & Generative AI Large Language Models (LLMs)Small Language Models (SLMs)Prompt EngineeringContext EngineeringRetrieval-Augmented Generation (RAG)EmbeddingsSemantic SearchAgentic AI PatternsMulti-Agent WorkflowsTool CallingFunction CallingModel EvaluationLLM Observability Model Engineering Fine-TuningSupervised Fine-TuningLoRAQLoRAQuantizationDistillationModel CompressionSynthetic Data GenerationModel BenchmarkingModel SelectionModel Routing Model Hosting & Serving Private LLM HostingOn-Prem Model DeploymentGPU-Based InferenceModel Serving APIsHigh-Availability InferenceAutoscalingLoad BalancingCachingBatch and Real-Time Inference AI Infrastructure & Frameworks KubernetesDockerKubeflowKServeRay ServeMLflowHugging FaceTransformersPyTorchPEFTDeepSpeedNVIDIA NIMTriton Inference ServerTensorRT-LLMvLLMTGISGLang Programming & Engineering PythonTypeScript or JavaScriptREST APIsMicroservicesCI/CDGitHub or Azure DevOpsAPI DesignDistributed SystemsCloud-Native EngineeringTest Automation Data & Knowledge Systems Vector DatabasesKnowledge GraphsDocument ProcessingMetadata ManagementData PipelinesObject StorageEnterprise SearchStructured and Unstructured Data Integration Roles & Responsibilities LLM Application Engineering Build enterprise-grade LLM-powered applications and intelligent agent capabilities. Design reusable LLM patterns, services, APIs, and accelerators. Develop model interaction patterns for:ReasoningSummarizationClassificationExtractionPlanningDecision supportBuild reusable prompt, context, retrieval, memory, and evaluation components. Support AI-for-SDLC agents across:RequirementsDesignCodingTestingSecurity ReviewDeploymentOperationsConvert AI use cases into scalable production solutions. Agent Factory Intelligence Layer Build core intelligence services for enterprise agents. Develop reusable capabilities for:PlanningTask decompositionReasoningTool usageAgent collaborationEnable agent-to-agent interaction and multi-agent orchestration. Integrate LLMs with:Agent runtimesTool registriesWorkflow enginesMCP-based gatewaysSupport human-in-the-loop, approval, escalation, and feedback workflows. Improve agent quality, accuracy, safety, and task completion. Prompt & Context Engineering Design reusable prompt engineering standards, templates, and libraries. Create:System promptsTask promptsRole promptsGuardrail promptsEvaluation promptsDevelop context engineering strategies for better grounding, relevance, and personalization. Optimize:Token usageContext windowsMemory injectionRetrieval inputsEstablish prompt versioning, testing, and governance practices. Retrieval-Augmented Generation (RAG) Design and implement enterprise RAG architectures. Build retrieval pipelines using:Enterprise documentsKnowledge repositoriesStructured dataMetadataOptimize:ChunkingEmbeddingsIndexingRankingRerankingRetrieval strategiesImprove grounding, citation quality, precision, recall, and factual accuracy. Build reusable retrieval services for agents and business domains. Partner with data and knowledge management teams to onboard trusted data sources. LLM / SLM Model Engineering Evaluate, build, fine-tune, deploy, and optimize LLMs and SLMs. Support domain-specific model development using approved datasets. Build supervised fine-tuning and model adaptation pipelines. Apply:LoRAQLoRADistillationQuantizationModel compressionEvaluate commercial, open-source, and internally hosted models. Select models based on:AccuracyLatencyCostData residencySecurityOperational requirements Private AI & On-Prem Hosting Build and support private AI capabilities for LLM/SLM hosting. Deploy models across:On-premisesHybridPrivate cloud environmentsSupport GPU-enabled model hosting. Optimize latency, throughput, concurrency, resiliency, and GPU utilization. Build secure inference endpoints for internal applications and agents. Support air-gapped and restricted AI environments. Partner with infrastructure and platform teams on private AI hosting. Model Serving & Inference Optimization Implement scalable model serving using modern inference frameworks. Build high-availability inference architectures. Optimize:Token throughputResponse latencyCost efficiencyInference performanceImplement:Model routingLoad balancingCachingFallback strategiesSupport batch and real-time inference. Develop reusable deployment templates for different model families. LLMOps / ModelOps / AgentOps Build operational practices for managing models and agents throughout their lifecycle. Implement observability for:PromptsRetrievalModel responsesLatencyCostFailuresDevelop evaluation pipelines for regression testing and continuous quality improvement. Monitor:Model driftResponse qualityHallucination indicatorsSafety risksSupport CI/CD and release management for:PromptsModelsAgentsRetrieval pipelinesBuild dashboards and metrics for AI quality, reliability, adoption, and operational readiness. AI Evaluation & Benchmarking Define and implement LLM evaluation frameworks. Measure:AccuracyGroundednessRelevanceHallucination rateToxicity riskSafety complianceTask completionUser satisfactionBuild automated test suites for prompts, agents, tools, and RAG pipelines. Benchmark models across enterprise use cases. Compare cloud, open-source, and on-prem models based on performance, cost, quality, and risk. Establish quality gates for production AI releases. Responsible AI, Security & Governance Implement Responsible AI controls in LLM applications and agent workflows. Develop guardrails for:Safe outputTool usageData accessEnterprise policy complianceSupport:Model risk managementAuditabilityTransparencyTraceabilityEnsure sensitive data is handled according to security and privacy requirements. Collaborate with Security, Enterprise Architecture, Risk, and Compliance teams. Support model and agent approval and production-readiness governance. Preferred / Additional Experience Experience deploying open-source models such as:LlamaMistralMixtralPhiGemmaQwenDeepSeekGraniteFalconDomain-specific modelsExperience with GPU infrastructure such as:NVIDIA H100H200B200B300A100L40SGH200AMD MI300XExperience with:Private AIHybrid AIAir-gapped AI environmentsExperience in regulated industries such as:HealthcareFinancial ServicesInsuranceExperience building:Enterprise copilotsAI assistantsAgent platformsExperience with:MCPTool registriesAgent runtimesEnterprise integration patternsExperience with Responsible AI, model governance, model risk management, and AI compliance. Experience optimizing AI workloads for:CostPerformanceLatencySecurity Role / Skills Summary Role: LLM EngineerEssential Skill: LLMPrimary Skill: AI and AutomationExperience: 8-10+ yearsWork Location: Remote USADuration: 1 yearCore Technologies: LLM, SLM, RAG, Generative AI, Agentic AI, Python, Kubernetes, MLflow, Hugging Face, PyTorch, vLLM, KServe, Docker, Vector Databases