Data Engineer, Sr.
Gemini is a leading global cryptocurrency and Web3 platform founded in 2014 by Cameron and Tyler Winklevoss. With a presence in over 70 countries, Gemini offers a comprehensive suite of simple, reliable, and secure crypto products and services designed for both individual and institutional clients. The company's mission is to unlock the next era of financial, creative, and personal freedom by providing trusted access to the decentralized future.
About the role
The Senior Data Engineer at Gemini will be a critical member of the data team responsible for designing, developing, and maintaining the data infrastructure that powers insights, reporting, analytics, and machine learning initiatives across the organization. This role involves making architectural decisions, mentoring junior engineers, and building high-scale, resilient data systems that deliver meaningful impact. The Senior Data Engineer will own the end-to-end delivery of data products within their domain, collaborating closely with cross-functional teams including product, analytics, machine learning, finance, operations, and engineering. The role requires a proactive approach to transforming, moving, and modeling data reliably while ensuring observability, resilience, and agility across all data workflows.
Responsibilities
- Design, develop, and maintain robust data infrastructure and pipelines for batch and real-time workloads
- Contribute to architectural decisions to optimize data systems for scalability and efficiency
- Build and manage scalable ETL/ELT pipelines using appropriate languages and frameworks
- Develop real-time data solutions such as CDC, streaming, and micro-batch processing for timely data delivery
- Collaborate with data scientists, ML engineers, analysts, and product teams to understand data requirements and define SLAs
- Establish data quality, validation, observability, and monitoring frameworks to ensure data integrity
- Investigate and troubleshoot complex production issues, including root cause analysis and performance bottlenecks
- Document data flows, architecture patterns, data dictionaries, and operational runbooks for transparency and knowledge sharing
Requirements
- 5+ years of experience in data engineering or similar roles
- Strong expertise in ETL/ELT pipeline design, implementation, and optimization
- Deep proficiency in Python and SQL for production-quality, maintainable, and testable code
- Experience with large-scale data warehouses such as Databricks, BigQuery, or Snowflake
- Solid understanding of software engineering fundamentals, data structures, and systems thinking
- Hands-on experience in data modeling including dimensional modeling, normalization, and schema design
- Experience building real-time or streaming data systems using Kafka, Kinesis, Flink, Spark Streaming, etc.
- Familiarity with CDC frameworks and data orchestration tools like Airflow
- Knowledge of data governance, lineage, metadata management, and data quality practices
Benefits
- Competitive starting salary
- Discretionary annual bonus
- Long-term incentive through equity grants for new hires
- Comprehensive health insurance plans
- 401(k) plan with company matching
- Paid parental leave
- Flexible time-off policies
- Hybrid work environment with options for remote work or in-office collaboration depending on location and role