QA Engineer (AI Systems)
Nexxa.ai · Sunnyvale, CA · 2 days ago
HybridQuality AssuranceFull-time
Key Responsibilities
- Design and build evaluation harnesses and regression suites for LLM-based agents, covering reasoning quality, tool-call correctness, task completion, and multi-turn coherence.
- Develop golden datasets and labeled test sets, including edge cases, ambiguous inputs, and adversarial prompts specific to industrial and operational contexts.
- Define and track quality metrics beyond simple accuracy — groundedness, hallucination rate, task success rate, latency/cost tradeoffs, and safety violations.
- Build automated pipelines that run evals on every model, prompt, or tool-integration change, and integrate them into CI/CD.
- Conduct structured red-teaming and adversarial testing (prompt injection, jailbreaks, tool misuse, unsafe actions) in partnership with security teams.
- Test agent behavior across the full action loop — planning, tool selection, tool execution, error recovery, and final output — not just the final response.
- Investigate and triage failures where the root cause could be the model, the prompt, the tool/API, or the orchestration logic.
- Partner with ML and backend engineers to translate eval failures into actionable, reproducible bug reports.
- Establish quality bars and sign-off criteria for new agent capabilities before they reach customer environments.
- Mentor other engineers on testing strategies specific to probabilistic, LLM-driven systems.
- Advocate for testability and observability in agent architecture from day one.
Qualifications
- 5+ years in QA/SDET roles, with demonstrated ownership of test strategy for complex systems.
- Hands-on experience testing LLM-based products, chatbots, or AI agents — you understand why traditional deterministic test assertions break down for generative systems.
- Practical experience with eval frameworks or tooling (e.g., promptfoo, DeepEval, RAGAS, LangSmith) or a track record of building your own.
- Strong scripting/programming ability (Python preferred) to build test automation, data pipelines, and eval tooling.
- Understanding of how LLM agents work: prompting, tool/function calling, context management, RAG, memory, and orchestration frameworks.
- Experience designing test data and labeled datasets, including sourcing, sampling, and managing dataset drift over time.
- Familiarity with LLM-specific failure modes: hallucination, prompt injection, context poisoning, tool misuse, goal drift, and non-determinism.
- Comfortable operating in ambiguity — defining what "correct" means for a task when there's no single right answer.
- Strong written communication skills for turning fuzzy quality signals into clear, actionable findings for engineering and product stakeholders.
PREFERRED EXPERIENCE
- With human-in-the-loop evaluation workflows (labeling pipelines, inter-rater reliability, rubric design).
- Background in ML/data science sufficient to read model evals and statistical significance.
- Experience red-teaming or doing adversarial/security testing on ML systems.
- Familiarity with observability/tracing tools for LLM applications (e.g., LangSmith, Arize, Langfuse, Weights & Biases).
- Experience testing AI systems in industrial, IoT, or operational technology (OT) environments.
- Prior experience setting up eval infrastructure from scratch at a startup or fast-moving team.