Senior Data Engineer (1043) - DataSF - Office of City Administrator
City and County of San Francisco · San Francisco, CA · 2 wk ago
Information TechnologyFull-time
About the role
The Senior Data Engineer will play a crucial role in designing, building, and maintaining the City's data infrastructure. They will be responsible for managing the central Snowflake data warehouse, developing scalable and resilient data pipelines, and championing best practices for data engineering.
Responsibilities
- Manage the central Snowflake data warehouse, including access control, security policies, resource monitoring, performance tuning, and cost optimization.
- Administer the platform with a focus on data democratization and accessibility, while protecting privacy and security.
- Build and maintain scalable and resilient pipelines to ingest and structure data from diverse sources.
- Use Terraform to define, deploy, and manage data infrastructure, ensuring reproducibility, version control, and production readiness.
- Champion and implement best practices for documentation, data modeling, warehouse architecture, SQL optimization, and testing.
- Collaborate with data scientists, analysts, product managers, software engineers, and nontechnical stakeholders to understand data requirements and build solutions that meet their needs.
- Proactively monitor the health of the data platform and pipelines, troubleshoot issues, and ensure high standards of data quality and availability.
Requirements
- Three (3) years of experience analyzing, installing, configuring, enhancing, and/or maintaining the components of a system or platform.
- Technical Knowledge: Hands-on experience administering and developing on managed cloud data platforms such as Snowflake, BigQuery, or Databricks.
- Demonstrated expertise in writing advanced, performant SQL, and using tools like dbt for SQL-based data transformation and modeling.
- Strong programming skills in Python (with libraries like pandas, PySpark) for data processing and automation.
- Proficiency with an Infrastructure as Code tool, with a preference for Terraform.
- Experience building and deploying data pipelines using orchestration tools like Azure Data Factory, Airflow, Dagster, or similar technologies.
- Deep understanding of data warehousing concepts, data modeling, and modern ELT principles.
- Understanding of data governance, data security, and data privacy principles.
- Experience with real-time data streaming technologies (e.g., Kafka, Kinesis, Snowpipe).
- Experience deploying and managing data pipelines for machine learning models.
Qualifications
- Education: An associate degree in computer science, computer engineering, information systems, or a closely related field from an accredited college or university OR its equivalent in terms of total course credits/units [i.e., at least sixty (60) semester or ninety (90) quarter credits/units with a minimum of twenty (20) semester or thirty (30) quarter credits/units in one of the fields above or a closely-related field].
- Experience: Three (3) years of experience analyzing, installing, configuring, enhancing, and/or maintaining the components of a system or platform.
- Substitution: One year of additional experience as described above may be substituted for the required degree.
- Completion of the 1010 Information Systems Trainee Program may be substituted for the required degree.