Backend AI Engineer
Nexxa.ai · Sunnyvale, CA · 2 days ago
HybridEngineeringFull-time
Key Responsibilities
- Design, build, and maintain backend services and APIs that power GenAI, LLM, and Computer Vision model integrations across Nexxa's products.
- Build and own core AI/ML infrastructure: model-serving pipelines, inference services, data pipelines, and embedding/vector stores.
- Architect scalable, production-grade systems for real-time and batch AI workloads across manufacturing, infrastructure, and logistics domains.
- Implement and optimize RAG systems, prompt/context pipelines, and orchestration layers connecting models to enterprise and operational data sources.
- Build robust APIs, microservices, and integration layers connecting AI systems to customer data, legacy systems, and existing infrastructure.
- Own the reliability, performance, and observability of backend AI systems — logging, monitoring, testing, and CI/CD for ML services.
- Collaborate closely with Forward Deployed Engineers, ML engineers, and product teams to translate customer and field requirements into reusable, hardened backend capabilities.
- Evaluate and integrate ML/CV/LLM models into production backend systems; manage model versioning, rollout, and deployment pipelines.
- Produce clear technical documentation: architecture diagrams, API specs, and runbooks for internal and customer-facing teams.
- Mentor engineers and contribute to internal backend engineering best practices.
Qualifications
- 4–8+ years of experience in backend software engineering, ML/platform engineering, or similar roles.
- Strong proficiency in TypeScript/Node.js (our primary backend language), with strong API and microservice design skills.
- Working proficiency in Python is a plus for ML/model integration work.
- Hands-on experience building and operating production backend systems at scale — distributed systems, databases, message queues.
- Experience integrating ML or Generative AI models (LLMs, multimodal models) into backend services — inference, orchestration, and evaluation.
- Solid understanding of cloud infrastructure (AWS, GCP, or Azure) and containerization (Docker, Kubernetes).
- Experience designing and operating data pipelines (batch and/or streaming) across structured and unstructured data.
- Hands-on experience building retrieval-augmented generation (RAG) systems and AI memory architectures — retrieval pipelines, vector stores, context management, and long-term/session memory for LLM applications.
- Strong grasp of system design fundamentals: scalability, reliability, security, and observability.
- Comfortable working cross-functionally with ML engineers, product, and customer-facing teams.
- Bachelor's degree (or higher) in Computer Science or a related field.
- PREFERRED: Familiarity with ML frameworks (PyTorch, TensorFlow, OpenCV) sufficient to integrate, serve, or evaluate models, even without training them yourself.
- Experience with MLOps tooling: model registries, feature stores, CI/CD for ML, and monitoring/observability for ML systems.
- Background in event-driven or real-time systems (Kafka, gRPC, WebSockets).
- Experience in industrial, IoT, or operational technology (OT) environments.
- Experience in startup or high-growth environments.