Senior Software Engineer - API & Data Connectors - MarTech/AdTech
Three Pillars Recruiting · San Francisco, CA · 3 wk ago
On-siteInformation TechnologyFull-time
This role sits in Data Products, which leverages the company's go-to-market, operations, and SRE functions and adheres to the company's engineering standards and tooling, while running its own product-focused, sprint-based development cycle. It’s the right fit for engineers who understand hyperscaling startups: rapid iteration, resourcefulness, and comfort with shifting priorities.
Responsibilities
- Architect Connector Abstractions: Design, extend, and maintain core Abstract Base Classes (Connector and WatcherConnector ABCs) that establish clean, typed SDK-like interfaces for all data ingestion.
- Build High-Performance Native Connectors: Develop zero-copy, memory-efficient native connectors for local filesystems, Amazon S3, S3 file watchers, and Apache Iceberg utilizing DuckDB relations, Apache Arrow, and Flowbee DataRef primitives.
- Own the Flask Runtime Service: Build, maintain, and optimize the Flask connector service that exposes standardized runtime operational contracts: catalog, check, discover, and read.
- Integrate Broad Source Catalogs: Maintain and scale our optional PyAirbyte integration (airbyte>=0.20.0), using PyAirbyte as an optional adapter backend to instantly unlock 300+ external Airbyte data sources.
- Optimize Data Transfer Semantics: Ensure low latency and high memory efficiency when materializing and passing data references (DataRef) through pipeline execution layers.
- Developer Experience & Extensibility: Ensure internal and external developers can easily author, test, and deploy custom connectors against our framework with minimal friction and strong typing guarantees.
Requirements
- 5+ years writing production software, with deep mastery of Python 3.10+ (async/await, strict type hints, ABCs, Pydantic, and clean OOP abstraction patterns).
- Experience building and operating microservices/APIs using Flask or similar lightweight web frameworks, with a focus on high-throughput JSON/stream serialization.
- Proven track record of building extensible SDKs, abstract interfaces, or pluggable framework architectures used by other engineers.
- Deep experience with modern embedded analytical engines and columnar formats: DuckDB relations, Apache Arrow, and Apache Iceberg.
- Hands-on experience with cloud storage ingestion patterns (Amazon S3, event-driven S3 file watchers, local storage abstractions).
- Familiarity with standard ELT/ETL lifecycle concepts (catalog, schema discovery, health check, read/stream).
- Experience integrating or extending data connector frameworks like PyAirbyte (airbyte>=0.20.0), Singer, or Meltano.
- Experience handling complex Python dependency management and packaging (Poetry, pip-tools) for modular framework plugins.
- Daily comfort with Linux, Docker, Git, automated testing (pytest), and CI/CD pipelines.
Nice to Have
- Familiarity with custom event-driven file watching and streaming architectures.
- Performance optimization experience dropping into Rust or C++ for bottleneck operations.
- Experience working with low-overhead data reference passing engines (e.g., Flowbee DataRef).
About the Role
The role is right for you if you appreciate elegant code abstractions just as much as raw data-throughput performance and enjoy building the foundational tools, framework ABCs, and SDKs that empower an entire data platform to talk to any data source on earth.