AI Systems Engineer - Data & State Management - Senior
About the role
We are seeking an AI Systems Engineer to own the stateful backbone of EY’s AI-native platform, including the data stores, memory tiers, and event streaming systems that hold and move every piece of state in EY’s Hybrid AI Multi-Environment Runtime (HAI).
Responsibilities
- Own the memory and data stores: relational and durable state (PostgreSQL, DBOS durable workflows), caching (Redis/Valkey), vector stores (Qdrant/Milvus/PGVector), knowledge graphs (Neo4j), and object/block storage (MinIO, OpenEBS Mayastor), across every environment and tenant.
- Own event streaming and async messaging: Apache Kafka (Strimzi), NATS JetStream (agent-to-agent), Debezium (change data capture), Apache Flink (stream processing), and Apicurio/CloudEvents (schema and event contracts).
- Build and operate streaming and CDC pipelines that move data reliably between stores and services, with schema governance and evolution that prevents breaking changes across producers and consumers.
- Make state multi-tenant and portable, ensuring isolation, performance, and consistent semantics whether running on managed cloud services or self-hosted OSS in an air-gapped environment.
- Provide the data and lineage substrate that downstream governance, observability, and AI knowledge capabilities depend on, as well as integrating with lineage tooling.
Requirements
- Deep expertise operating production databases and data stores at scale, including relational, key-value, vector, graph, and object storage.
- Strong command of streaming and event-driven architectures (Kafka, NATS, CDC, stream processing) and the consistency tradeoffs they involve.
- A durability-first mindset: thinking in terms of consistency, recoverability, blast radius, and data correctness under failure.
- Ability to operate stateful systems consistently across managed cloud and self-hosted OSS in cloud, on-prem, edge, and air-gapped environments.
- Strong grasp of schema governance and evolution, preventing breaking changes across producers and consumers.
- Strong communicator able to guide consuming teams toward the right storage and streaming patterns.
- Orientation toward reliability and toil reduction through automation and infrastructure-as-code for data systems.
Attributes for Success
- Hands-on expertise with relational databases (PostgreSQL) and caching (Redis/Valkey), including HA, replication, and backup/recovery.
- Exposure to AI/ML data patterns — embeddings, retrieval, feature/state stores for agentic workloads.
- Production experience with vector and/or graph databases (Qdrant, Milvus, PGVector, Neo4j) in AI/ML contexts.
- Experience with stream processing (Apache Flink) and event/schema governance (CloudEvents, schema registry).
- Experience running stateful systems on Kubernetes (operators, persistent volumes, object/block storage such as MinIO/OpenEBS).
- Experience defining clean ownership boundaries and data/schema contracts with platform, trust, runtime, and delivery teams.
- Experience with durable workflow engines (DBOS, Temporal, or equivalents).
- Experience with data lineage and metadata tooling (OpenLineage, Marquez, or equivalents).
- Experience operating data systems under compliance, security, or regulatory constraints, including data-at-rest and in-transit protection.
- Experience with multi-tenant data isolation and per-tenant performance management.
- Experience with cross-environment replication and DR strategies tiered by RPO/RTO (e.g., MirrorMaker, CloudNativePG PITR).
- Exposure to regulated delivery environments (financial services, tax, healthcare, risk).
Pay
The base salary range for this job in all geographic locations in the US is $106,900 to $176,500. The base salary range for New York City Metro Area, Washington State and California (excluding Sacramento) is $128,400 to $200,600. Individual salaries within those ranges are determined through a wide variety of factors including but not limited to education, experience, knowledge, skills and geography.
Schedule
Our expectation is for most people in external, client serving roles to work together in person 40-60% of the time over the course of an engagement, project or year.