Senior Software Engineer, Data
ICON · Austin, TX · 2 days ago
On-siteEngineeringFull-time
About the role
In this role, you will be responsible for designing, building, and maintaining the data infrastructure that powers our data-driven products and services.
Data engineering responsibilities
- Design and build robust data pipelines to ingest, transform, and load data from a variety of internal and external sources.
- Develop and maintain large, high quality datasets that power machine learning models and analytical products.
- Build scalable data storage, processing, and serving infrastructure on cloud platforms.
- Develop tooling and services for data labeling, data review, and dataset curation at scale.
- Find and evaluate new external data sources - manage relationships with data partners and vendors.
- Contribute to CI/CD practices and engineering standards across the data platform.
Data science responsibilities
- Partner with ML engineers and researchers to design feature pipelines and experiment infrastructure.
- Apply statistical and analytical techniques to assess dataset quality, identify gaps, and surface insights.
- Develop efficient algorithms and data models to curate data and maintain high quality and consistency.
- Design and evaluate metrics to measure dataset health and model readiness.
- Translate ambiguous research questions into well defined data problems with measurable outcomes.
Minimum qualifications
- 7+ years of software development experience.
- 5+ years of hands on experience managing and engineering data at scale.
- Proficiency in Python, TypeScript and SQL.
- Experience with cloud data services and storage technologies (e.g., S3, Redshift, BigQuery, Snowflake).
- Familiarity with statistical analysis and core data science concepts (feature engineering, data distributions, model evaluation etc).
- Degree in Computer Science, Statistics, a related technical field, or equivalent experience.
- Strong problem solving skills with the ability to work independently and drive projects end-to-end.
Preferred skills and experience
- Experience building large datasets for machine learning training and evaluation.
- Proficiency with data science libraries (pandas, NumPy etc.).
- AWS experience, including CDK or similar managed services.
- Experience with workflow orchestration tools (Airflow, Prefect etc).
- Experience with modern CI/CD workflows
- Experience partnering with external organizations or managing 3rd party data vendors.