Lead Data Engineer – Vice President
Citi · Jersey City, NJ · 4 wk ago
Hybrid$142k–$213k/yrFull-time
About the role
Citi is looking for a Lead Data Engineer – AI & Distributed Analytics to architect and scale the enterprise data platforms that power our advanced analytics, machine learning, and Generative AI capabilities.
Responsibilities
- Define and drive the technical vision, architecture, and roadmap for Citi's enterprise Data Lake and Analytics platform, ensuring alignment with cloud-native and hybrid-cloud strategies.
- Lead, mentor, and develop a team of data engineers, setting technical standards and fostering a culture of continuous learning and agile delivery.
- Design and build robust batch and real-time streaming data pipelines using technologies such as Apache Iceberg, Starburst, Startree/Apache Pinot, Apache Kafka, and Apache Flink to support enterprise-scale data processing.
- Architect and implement Feature Stores to standardize feature engineering and ensure consistent, reliable data delivery for model training and real-time ML inference.
- Optimize data storage and query performance across Enterprise Data Lakes and Lakehouses, leveraging platforms such as Delta Lake, Apache Iceberg, and Databricks.
- Implement end-to-end metadata management and data lineage tracking to provide full visibility into how data flows from source systems through to AI models and analytics consumers.
- Evaluate and prototype emerging data technologies and frameworks, translating findings into actionable recommendations that keep Citi's data platform at the cutting edge.
Requirements
- Bachelor's degree, university degree, or equivalent professional experience in Computer Science, Data Engineering, Information Systems, or a quantitative field.
- 6 or more years of professional experience in data engineering, software engineering, or data platform development, including 3 or more years in a technical lead or engineering leadership role.
- Expert-level proficiency in Python and SQL, with hands-on experience designing and delivering large-scale, distributed data systems.
- Deep expertise in big data frameworks including Apache Iceberg, Startree/Apache Pinot, and Hive/HDFS, with a strong command of distributed computing principles.
- Hands-on experience building and operating real-time streaming pipelines using Apache Kafka and Apache Flink at high volume and scale.
- Demonstrated ability to design and build data infrastructure specifically supporting machine learning, advanced analytics, or AI applications in production environments.
- Experience working with modern table formats and transactional storage layers such as Delta Lake, Apache Iceberg, or Apache Hudi.
- Solid experience with workflow orchestration tools such as Apache Airflow or equivalent modern orchestrators to manage complex pipeline dependencies.
Skills
- Bachelor's degree in Computer Science, Data Engineering, Information Systems, or a related quantitative discipline.
- Additional programming proficiency in Java.
- Experience with cloud data platforms such as Databricks, including cluster tuning and lakehouse architecture optimization.
- Familiarity with vector databases — such as Milvus, Pinecone, Qdrant, or Chroma — for GenAI and large language model use cases.
- Understanding of dimensional modeling, Data Vault, and schema-on-read/write design patterns for enterprise analytics workloads.
- Hands-on experience with containerization and infrastructure tooling including Docker, Kubernetes, and Terraform for Infrastructure as Code.