Jobs · Engineering · California

Principal AI Engineer - LLM/Agents

SB Energy · Redwood City, CA · 2 days ago
On-siteEngineering$160k–$190k/yrFull-time

About the role

Join us at SB Energy, a leading infrastructure company backed by SoftBank Group and Ares Infrastructure, where we combine cutting-edge innovation with best-in-class execution to deliver reliable and secure data center and power infrastructure.

Responsibilities

  • Define the technical architecture for SB Energy agent workflows, including ChatGPT, ROCstar, MCP, agents, skills, tools, APIs, retrieval, structured outputs, trace metadata, evals, and observability.
  • Create standards for skill design, prompt structure, tool descriptions, tool schemas, API response formats, source citation behavior, structured outputs, and workflow handoffs.
  • Build evaluation frameworks for agent workflows, including gold datasets, regression tests, numerical reconciliation checks, rubric-based grading, tool-call correctness checks, and human feedback loops.
  • Lead the design of observability for agents and tools, including workflow logs, cost, latency, token usage, tool success rate, bad-answer rate, eval pass rate, user acceptance, and incident tracking.
  • Partner with domain Forward Deployed Engineers to convert high-value workflows into measurable agent systems with explicit inputs, outputs, owners, permissions, evals, runbooks, release gates, and continuous improvement loops.
  • Partner with the Enterprise Systems and Agent Platform role on MCP servers, connector reliability, RBAC, secrets, audit logging, deployment patterns, API governance, and production platform readiness.
  • Review agent workflow designs before production release and define go/no-go criteria for quality, safety, reliability, cost, latency, security, and operational support.
  • Create reusable templates for agent specifications, skill specifications, eval protocols, workflow scorecards, incident reviews, production readiness checklists, and workflow-level success metrics.
  • Mentor FDEs, analysts, engineers, and implementation partners on eval-driven development, tool-based workflow design, observable agent operations, and production-quality AI systems.
  • Identify recurring failure modes across agents and tools and turn them into tests, standards, instrumentation, documentation, and platform improvements.

Qualifications/Requirements

  • Bachelor's degree in Computer Science, Software Engineering, Data Science, Machine Learning, Information Systems, or a related technical field required.
  • Master's degree in Computer Science, AI/ML, Data Systems, Software Engineering, or a related field preferred.
  • 10+ years of experience in software engineering, data science, machine learning systems, enterprise AI, platform engineering, or decision systems, with demonstrated technical leadership across complex production workflows.
  • Experience designing, building, or governing agentic AI systems, LLM applications, tool-calling workflows, evaluation systems, observability systems, or production AI platforms.
  • Strong understanding of modern LLM workflow design, including prompts, skills, tools, retrieval, structured outputs, APIs, orchestration, context management, token economics, and human-in-the-loop controls.
  • Strong engineering judgment around production readiness, failure modes, logging, testing, monitoring, cost, latency, security, permissions, and operational support.
  • Experience building eval datasets, regression tests, model or agent scorecards, acceptance criteria, telemetry dashboards, release gates, and production quality standards.
  • Experience with Python, APIs, data pipelines, cloud platforms, CI/CD, logging and monitoring systems, analytics workflows, and enterprise integrations.
  • Experience with OpenAI, ChatGPT Enterprise, Responses API-style workflows, MCP, Azure, Databricks, SharePoint, Teams, Outlook, Monday.com, or enterprise connectors preferred.
  • Ability to translate ambiguous business workflows into practical system designs, measurable success criteria, and reusable implementation patterns.
  • Ability to work across domain FDEs, engineers, analysts, business leaders, IT, security, legal, compliance, and external implementation partners.

Location

Location options include Houston, TX, San Francisco Bay Area, CA, Denver, CO, or remote option available. The position may require up to 10% domestic travel.

Pay

Base Pay - $160000-$190000. The pay range mentioned above is a guideline. We tailor each offer based on your unique skills, experience, location, and market benchmarks—while ensuring internal parity across our teams. The compensation range provided reflects base pay. Total rewards may also include performance-based bonuses, equity grants, and a market leading comprehensive health and wellness benefits package. Final details will be discussed during the later stages of the hiring process.

Similar jobs