AI Engineer
ClearpointCo · Houston, TX · Yesterday
Engineering$120k–$125k/yrFull-time
Duties
- Architects and implements on-premises AI platform from the ground up, leveraging newly acquired on-prem GPU hardware (H200-based) to deliver enterprise AI capability with no cloud dependency.
- Designs and builds LLM-based applications, including RAG pipelines, agentic search, and vector retrieval systems, using self-hosted frameworks such as Haystack, LlamaIndex, Ollama, or comparable tools.
- Establishes foundational standards, patterns, and best practices for AI development, since none currently exist.
- Deploys, tunes, and monitors models on on-prem GPU infrastructure, optimizing for throughput, latency, and resource utilization across shared workloads.
- Collaborates with Data Engineering to define data contracts, feature pipelines, and integration points between the MS SQL Server/DB2 data warehouse and AI systems.
- Evaluates and prototypes emerging AI tooling, both open-source and commercial, for fit within a strict on-premises, no-cloud-egress environment.
- Develops internal tools and APIs that expose AI capabilities to other departments, such as virtual agent support, document processing, and ticket triage.
- Performs other job related duties as assigned.
Requirements
- 3+ years of experience building and deploying machine learning or LLM-based applications in production, including experience architecting systems rather than solely implementing within existing ones.
- Work involves regular collaboration with Data Engineering, departmental leadership, and end users across the credit union.
- Ability to clearly explain complex technical concepts to non-technical stakeholders and to operate with a high degree of independence and sound judgment is essential.
- Strong Python skills, including experience with ML/AI frameworks (PyTorch, Hugging Face Transformers, LangChain/LlamaIndex/ Haystack, or similar).
- Experience with vector databases and embedding-based retrieval systems.
- Hands-on experience with containerized deployment (Docker) on Linux (Debian/Ubuntu) servers.
- Demonstrated ability to operate independently in greenfield environments, building from zero with limited existing infrastructure or precedent.
- Understanding of data privacy and PII handling practices.
- Experience operating self-hosted LLMs (Ollama, vLLM, text-generation-inference) on GPU hardware, including GPU resource planning and capacity management, preferred.
- Must have good communication skills.
- Able to maintain a high level of confidentiality at all times.