Jobs · Information Technology · New York

Staff Data Engineer

hackajob · New York, NY · 2 mo ago
On-siteInformation Technology$138/hrFull-time

About the Company

Twenty Technologies, founded in 2024, industrializes offensive cyber operations for the U.S. and its allies. Based in Arlington, Virginia, Twenty has raised $138M from leading investors.

Role Summary

You will own the data infrastructure that powers Twenty’s cyber operations applications and capabilities. This role involves building a durable, high-performance data lake and the pipelines, schemas, and query patterns that make petabyte-scale datasets usable and economical. You’ll partner closely with engineers and intelligence analysts to transform messy, high-volume operational data into reliable, well-modeled systems that drive real missions. You’ll also lead technical initiatives and mentor other engineers as we scale what we can support and ship.

Who You Are

  • Think in systems: data modeling, storage formats, compute engines, and access patterns all have to fit together.
  • Opinionated about schema and index design, and you can explain tradeoffs clearly.
  • Default to measurable reliability: data quality, lineage, repeatability, and operational excellence.
  • Comfortable working with ambiguous datasets and evolving requirements without lowering standards.
  • Collaborate tightly across roles, especially with engineers and analysts who need fast, correct answers.
  • Take leadership seriously—mentoring others, raising the bar, and driving initiatives to completion.
  • Motivated by national security outcomes and want your work to matter in the real world.

What You'll Do

  • Lead the development and operation of a data lake for cyber operations and intelligence data.
  • Design schemas, partitions, and indexes that make complex datasets performant and cost-effective to query.
  • Partner with engineers and intelligence analysts to define query patterns and data products for mission use cases.
  • Build and evolve ETL pipelines that are observable, recoverable, and resilient to upstream change.
  • Drive technical initiatives end-to-end, from architecture decisions through production rollout and iteration.
  • Establish best practices for data quality, documentation, and operational ownership across the platform.
  • Mentor engineers on data modeling, performance tuning, and production-grade pipeline design.
  • Identify bottlenecks in storage/compute/query layers and ship improvements with clear performance wins.

Must Have

  • 8+ years of experience in data engineering and/or data architecture.
  • Mastery-level expertise building ETL pipelines and operating them in production.
  • Deep experience with data lake architecture and systems used to query data lakes.
  • Strong schema and index design skills, including partitioning, indexing, and clustering strategies.
  • Experience with column-oriented databases in production environments.
  • Built data systems from scratch (not just maintained existing platforms).
  • Proven leadership experience mentoring engineers and driving technical initiatives.
  • U.S. citizenship and meeting the role’s security requirements.

Nice To Have

  • Experience with key-value datastores.
  • Experience with streaming and message queue systems.
  • Experience with graph database technologies.
  • Experience supporting analysts or operational users with high-stakes data needs.

Tech Environment (You Might Work With)

  • Data lakes: Apache Iceberg, Delta Lake, Apache Hive
  • Query engines: Trino, Presto, AWS Athena, Apache Spark
  • Column stores: ClickHouse, Amazon Redshift, Google BigQuery
  • ETL / orchestration: Airflow, AWS Glue, NiFi, ClickPipe
  • Streaming / queues: Kafka, RabbitMQ, NATS, AWS Kinesis
  • Graph: Neo4j, AWS Neptune, Memgraph, Apache AGE

Benefits

  • Health: Medical, dental, and vision plan options. Life / AD&D, disability coverage options.
  • Family: Paid parental leave for eligible full-time employees. 12 weeks for birthing parents, 4 for non-birthing parents, 6 weeks for adoptive, foster, or intended parents through surrogacy.
  • Vacation: Paid holidays and flexible PTO. Take what you need.
  • Retirement: 401(k) with pre-tax and Roth options. HSA/FSA options, dependent care FSA.
  • At the office: Commuter benefits. On-site garage parking. Bike storage. Building fitness center. Desk setup stipend.

Benefits Vary by Location, Role, and Eligibility

Full plan details provided during the interview and offer process.

Qualifications

  • You have 8+ years of experience in data engineering and/or data architecture.
  • You have mastery-level expertise building ETL pipelines and operating them in production.
  • You have deep experience with data lake architecture and systems used to query data lakes.
  • You have strong schema and index design skills, including partitioning, indexing, and clustering strategies.
  • You have experience with column-oriented databases in production environments.
  • You have built data systems from scratch (not just maintained existing platforms).
  • You have proven leadership experience mentoring engineers and driving technical initiatives.
  • You are a U.S. citizen and can meet the role’s security requirements.

Pay

Details TBD

Schedule

Details TBD

Similar jobs

Staff Data Engineer

Necessary VenturesBerkeley, CA· 1 mo ago
Information Technologyapply on copperhome.com

Staff Data Engineer

HCA HealthcareNashville, TN· 2 mo ago
Healthcareapply on careers.hcahealthcare.com

Staff Data Engineer

Awetomaton LtdSt Louis, MO· 2 mo ago
Information Technology$25k/yrapply on grnh.se

Staff Data Engineer

Self Financial, Inc.Austin, TX· 2 mo ago
Information Technology$160k–$195k/yrapply on grnh.se

Staff Data Engineer

ImprintBuffalo-Niagara Falls Area· 2 mo ago
Information Technologyapply on jobs.ashbyhq.com

Staff Data Engineer

XometryNorth Bethesda, MD· 2 mo ago
Information Technology$180k–$200k/yrapply on job-boards.greenhouse.io