Staff Data Engineer - Credit Karma
Intuit · San Diego, CA · 1 mo ago
On-siteInformation Technology$203k–$274k/yrFull-time
Responsibilities
- Design, build, and maintain high-throughput, low-latency data frameworks used across Credit Karma's engineering organization, including ETL templates, persistence libraries, and streaming data pipelines
- Develop and extend Scala-based microservices and frameworks built on Finagle, Akka Streams, and gRPC that process petabytes of data daily
- Build and optimize cloud-native data pipelines on Google Cloud Platform using Dataflow (Apache Beam), Flink, Spark, Pub/Sub, BigQuery, and Spanner
- Own and evolve our Kafka-based streaming infrastructure — designing producers, consumers, and connectors that handle hundreds of terabytes of events per day with strict latency and durability guarantees
- Build persistence frameworks that provide a unified, type-safe API for reading and writing across Spanner, MySQL, and BigQuery
- Design and implement encryption, decryption, and fine-grained access control capabilities as reusable framework features, ensuring compliance with data governance requirements
- Create self-service developer tooling — CLI tools, templates, and onboarding automation — that reduces the time for other teams to adopt the data platform from weeks to hours
- Drive technical design through architecture reviews and Technical Design Documents (TDDs), influencing decisions across the broader Data organization
- Participate in on-call rotations and build observability (dashboards, alerting, metrics) into every system you ship
Requirements
- 7+ years of professional software engineering experience building backend services and data infrastructure in Scala, Java, or a similar JVM language
- 7+ years of experience designing and operating high-throughput, low-latency distributed systems that process data at petabyte scale
- 3+ years of experience with streaming and messaging platforms such as Apache Kafka, Google Pub/Sub, or equivalent
- 3+ years of experience building data pipelines on a major cloud platform (GCP, AWS, or Azure), including services like Dataflow, BigQuery, Spanner, or their equivalents
- Professional experience with RPC frameworks such as Finagle, gRPC, or Akka for building production-grade service-to-service communication
- Strong understanding of software engineering best practices including CI/CD, version control (Git), code review, and automated testing
- Experience building reusable frameworks, SDKs, or platform libraries consumed by other engineering teams — you think about developer experience as a product
- Experience with Apache Beam (Dataflow) including custom transforms, side inputs, windowing strategies, and pipeline optimization
- Experience with Change Data Capture (CDC) patterns, particularly MySQL binlog-based replication to analytical stores
- Experience with data encryption at rest and in transit, including key management (KMS/GSM), SPIFFE/mTLS, and certificate authority integration
- Experience with schema management, data governance, data catalog, data quality and lineage management frameworks in a large-scale production environment
- Knowledge of AI/ML and GenAI technologies — LLMs, RAG, Semantic Search (e.g Vertex AI search, AWS cloud search) and Knowledge Graph (e.g neo4j, ..)
- Familiarity with infrastructure-as-code, Kubernetes (GKE), and container-based deployment models
- Track record of mentoring engineers and driving technical alignment across teams through design documents and architecture reviews
Qualifications
- Professional experience with Scala, Java, or a similar JVM language
- Experience with high-throughput, low-latency distributed systems processing data at petabyte scale
- Experience with streaming and messaging platforms such as Apache Kafka, Google Pub/Sub, or equivalent
- Experience with cloud-native data pipelines on Google Cloud Platform using Dataflow (Apache Beam), Flink, Spark, Pub/Sub, BigQuery, and Spanner
- Experience with RPC frameworks such as Finagle, gRPC, or Akka for building production-grade service-to-service communication
- Strong understanding of software engineering best practices including CI/CD, version control (Git), code review, and automated testing
- Experience building reusable frameworks, SDKs, or platform libraries consumed by other engineering teams — you think about developer experience as a product
- Experience with Apache Beam (Dataflow) including custom transforms, side inputs, windowing strategies, and pipeline optimization
- Experience with Change Data Capture (CDC) patterns, particularly MySQL binlog-based replication to analytical stores
- Experience with data encryption at rest and in transit, including key management (KMS/GSM), SPIFFE/mTLS, and certificate authority integration
- Experience with schema management, data governance, data catalog, data quality and lineage management frameworks in a large-scale production environment
- Knowledge of AI/ML and GenAI technologies — LLMs, RAG, Semantic Search (e.g Vertex AI search, AWS cloud search) and Knowledge Graph (e.g neo4j, ..)
- Familiarity with infrastructure-as-code, Kubernetes (GKE), and container-based deployment models
- Track record of mentoring engineers and driving technical alignment across teams through design documents and architecture reviews
Pay offered is based on factors such as job-related knowledge, skills, experience, and work location. To drive ongoing fair pay for employees, Intuit conducts regular comparisons across categories of ethnicity and gender. The expected base pay range for this position is: $202,500 - $274,000 depending on location.