Data Engineer - Onboarding
Sardine · California, United States · 4 wk ago
RemoteRemoteInformation TechnologyFull-time
About the role
We are looking for a Senior Data/ML Engineer to own the data and machine learning foundation that Sardine's compliance decisions run on. This role involves:
- Designing and implementing streaming and batch pipelines for data ingestion.
- Building and evolving the feature platform for computing features in real-time and batch.
- Maintaining feature correctness through reconciliation, recomputation tests, and drift monitoring.
- Productionizing fraud and identity ML models using Vertex AI and Kubeflow.
- Engineering KYC, AML, and identity risk signals, integrating and hardening new data sources.
- Ensuring the platform is safe by construction, including field-level encryption and data residency enforcement.
- Setting technical direction and raising the team's ceiling through design documentation, reviews, mentoring, and deciding what to build versus buy.
Responsibilities
You will write production code, make architectural calls that outlive your tenure, and raise the bar for how a small team ships fraud ML. You will work directly with data scientists, backend engineers, and the fraud analysts who use what you build.
Requirements
- 8+ years building production data and ML systems, with real ownership of both the pipeline side and the model side.
- Deep Python and strong SQL skills.
- Hands-on experience with a modern cloud data stack: GCP strongly preferred (BigQuery, Dataflow, Dataproc, Pub/Sub, Bigtable, Composer, Vertex AI) or the AWS equivalents, plus Docker, Kubernetes, Terraform, and CI/CD.
- Practical ML engineering depth: feature stores and feature pipelines, training/serving skew, gradient-boosted tree models, class imbalance and rare-event modeling, threshold and cost-sensitive tuning, model monitoring and drift detection, and explainability.
- Experience with high-volume, low-latency serving where a feature fetch has a few hundred milliseconds and there is no retry budget.
- Domain experience in fraud, risk, payments, lending, or identity/KYC — or the demonstrated ability to get fluent in a regulated domain fast.
- Comfort with data governance in a regulated environment: PII, encryption, access control, regional data residency, auditability.
- Strong written communication. You can explain a modeling tradeoff to a fraud analyst and a pipeline design to a backend engineer, and you write things down.
- A bias toward action and comfort in ambiguity. Much of this role is deciding what should exist, then building it.
Qualifications
- Experience supporting customer-facing ML — bring-your-own-model integrations, model explainability for adverse action or regulatory review, or shadow/challenger scoring frameworks.
- Experience in high-growth B2B SaaS, or as an early data/ML hire who built the function rather than inherited it.
Skills
- Experience with distributed processing frameworks (Spark, Beam, or Flink).
- Understanding of streaming semantics (windowing, watermarks, late data, exactly-once versus at-least-once, and where correctness actually breaks).
- Knowledge of high-volume, low-latency serving environments.
- Experience with fraud, risk, payments, lending, or identity/KYC domains.
Benefits
- Generous compensation in cash and equity.
- Early exercise for all options, including pre-vested.
- Work from anywhere: Remote-first Culture.
- Flexible paid time off and Year-end break.
- Health insurance, dental, and vision coverage for employees and dependents - US and Canada specific.
- 4% matching in 401k / RRSP - US and Canada specific.
- MacBook Pro delivered to your door.
- One-time stipend to set up a home office — desk, chair, screen, etc.
- Monthly meal stipend.
- Monthly social meet-up stipend.
- Annual health and wellness stipend.
- Annual Learning stipend.