Jobs · Information Technology · California

Senior AI Platform & Agentic Infrastructure Engineer

OKX · San Jose, CA · 1 mo ago
Information Technology$178k–$321k/yrFull-time

About the role

OKX’s Internal Audit function has an early but working AI-native capability: a multi-agent platform (Hive Mind), agentic workflows, data pipelines, and AI-enabled tools that let a very small team punch far above its weight. Your job is not to rebuild it at today’s maturity. Your job is to take it to a regulator-grade production standard and well beyond, and to drive AI Enablement across the department.

Responsibilities

  • Inherit, operate, and progressively migrate the working prototype estate (Google Workspace–native automation across Apps Script, Drive, and the Docs/Sheets/Slides APIs; locally scheduled jobs; and Claude Code agent tooling) to the target platform without interrupting daily and board-cycle workflows.
  • Reuse and integrate with OKX's prevailing and evolving AI capabilities and data infrastructure, including enterprise-approved models and gateways, Model Context Protocol (MCP) servers and connectors, security tooling, and group data platforms, before building parallel capability.
  • Re-architect Hive Mind into a resilient AWS or GCP platform with high availability (HA), disaster recovery (DR), defined service-level objectives (SLOs), and full observability, and own it in production.
  • Build the agentic runtime and harness: orchestration, multi-model routing that sends each task to the right model, tool, MCP, and plugin integration across model providers, and the evaluation, red-team, and regression harnesses that grade agents before and after production, with evaluation gates, judge calibration, and cost controls that hold for each model, including prompt-injection and data-exfiltration threat modeling for agents that read untrusted content.
  • Build the Responsible-AI and model-governance layer: hallucination, bias, and drift controls, output validation, guardrails, and complete logging. Extend the existing governance design (human-calibrated judge gates, golden-set regression, provenance registry) rather than replacing it. New model providers enter through the same governance and approved-tooling review, not around it.
  • Engineer data infrastructure with provenance and lineage: immutable audit trails, versioned evidence, reproducible pipelines, and traceability from source to report.
  • Stand up the computer-assisted audit technique (CAAT) and continuous-monitoring data foundation: analytics over full populations, with exceptions streamed in real time. Auditors and the hiring manager define the audit logic; you make it run at production grade.
  • Own the cloud foundation: infrastructure-as-code, CI/CD, identity, networking, observability, and cost controls, and make the platform examinable.

Requirements

7+ years building and operating resilient backend or platform systems in production, including on-call ownership over time.

Proven brownfield migrations: you have taken a founder-built or prototype system to production grade while it stayed in daily use.

Strong engineering fundamentals: data structures and algorithms, fluent Python, strong SQL and data modeling, plus one additional systems language (TypeScript/Node or Go), with the full-stack range to connect the infrastructure yourself.

Agentic runtime and harness engineering, proven by building: the field is too young to demand years of it, so we weigh real systems shipped over tenure. You have genuinely built model-agnostic agent orchestration, model routing and evaluation across providers, agent SDKs (Claude Agent SDK or equivalent), MCP servers, skill- and hook-based agent tooling, and the evaluation and red-team harnesses that grade agent behavior, with responsible-AI controls (hallucination, bias, drift).

Data engineering with provenance and lineage, and security and data-protection engineering by default (encryption, identity and access management, secrets, retention).

You build systems that could withstand external audit, by design.

Cloud (AWS or GCP) plus resilience engineering: infrastructure-as-code (Terraform), CI/CD, HA, DR, SLOs, and observability, backed by automated testing and documentation.

Ship inside locked-down enterprise environments (TLS-intercepting proxies, endpoint detection and response tooling, restricted installs, security guardrails, OAuth admin consent) without treating security as someone else’s problem.

Partnership: you translate audit needs into systems, explain technical risk to non-engineers, and drive adoption.

Qualifications

Multi-agent orchestration frameworks, MCP servers, and tooling and plugin development across the Anthropic, OpenAI, and Google model ecosystems.

Model-risk or AI-governance program experience; frameworks such as the NIST AI Risk Management Framework.

Continuous-auditing or continuous-controls-monitoring platforms; streaming and real-time data at scale; statistical anomaly detection and applied machine learning beyond LLMs; vector stores and retrieval; LLM cost engineering.

Experience passing external audit, SOC 2, or SOX (helpful, not required).

Crypto and blockchain literacy; regulated financial-services, fintech, or crypto experience.

Skills

Competitive total compensation package

L&D programs and Education subsidy for employees' growth and development

Variety of team building programs and company events

Wellness and meal allowances

Comprehensive healthcare schemes for employees and dependants

Benefits

Notice

Pay

The salary range for this position is $178,000 - $321,000

Schedule

Salary offered depends on a variety of factors, including job-related knowledge, skills, experience, and market location.

Similar jobs