Sr AI ENGINEER - Software
Snowrelic Inc · Jersey City, NJ · 2 wk ago
EngineeringFull-time
About the role
The client is a technology organization operating at the scale of a major financial enterprise, building AI agents that serve millions of customers directly. This is not a chatbot bolt-on or an innovation-lab experiment: the agentic platform sits in the production path of real customer conversations, with the reliability, latency, and safety bar that implies.
Responsibilities
- Design, build, and maintain production-grade LLM-based agentic systems.
- Architect and implement robust, scalable solutions for real customer conversations.
- Work with Python (primary agent/ML stack) and Java (service integration) daily.
- Develop and optimize agent components using LangChain, LangGraph, prompt/context engineering, tool/function calling, and RAG pipelines.
- Fine-tune, debug, and reason about model behavior using PyTorch and model fundamentals.
- Manage LLM serving/inference using vLLM or comparable (TGI, TensorRT-LLM) and integrate commercial APIs (OpenAI, Anthropic, Bedrock/Vertex).
- Apply distributed-systems fundamentals: APIs, queues, caching, observability, and CI/CD.
Requirements
- 7+ years of software engineering experience, with senior/lead-level ownership of production systems.
- Proven hands-on delivery of LLM-based agentic systems in production (not coursework or POCs).
- Strong proficiency in Python and Java, both used daily in this role.
- Depth in the standard agent stack: LangChain, LangGraph, prompt/context engineering, tool/function calling, RAG pipelines.
- Working knowledge of PyTorch and model fundamentals for fine-tuning, debugging, and tradeoff analysis.
- Experience with LLM serving/inference (vLLM, TGI, TensorRT-LLM) and commercial APIs (OpenAI, Anthropic, Bedrock/Vertex).
- Solid distributed-systems fundamentals: APIs, queues, caching, observability, CI/CD.
Qualifications
Major Plus:
- Hands-on experience with agent-builder platforms such as Sierra or Decagon (deploying, extending, or evaluating for customer-service use cases).
- Voice agents, real-time/streaming inference, or contact-center integration.
- Experience in regulated industries (finance, healthcare) with model risk, compliance review, or audit trails.
- Eval frameworks (LangSmith, Braintrust, custom harnesses) and LLM safety/guardrail tooling.
Location: NY/NJ, Columbus, or Dallas (On-Site)