Associate
About the role
Drive the architectural design, development, and implementation of mission-critical enterprise big data operational platforms and data migration initiatives. Leverage cloud-native solutions to enable advanced data analysis and generative AI outcomes, driving optimization across the investment product sales lifecycle.
- Architect and execute comprehensive end-to-end platform migration strategies for legacy Apache Hadoop/Apache Hive distributed data processing systems to modern cloud-native architectures utilizing Databricks Lakehouse Platform, Apache Airflow orchestration, and Snowflake cloud data warehouse.
- Design and implement comprehensive self-service data pipeline frameworks having complex data transformation logic converting MapReduce paradigms and Hive Query Language (HiveQL) operations to Apache Spark 3.x DataFrame API and PySpark distributed computing frameworks.
- Develop automated, event-driven and optimized Apache Airflow Directed Acyclic Graph (DAG) workflows supporting diverse data ingestion and transformation of structured and semi-structured formats.
- Integrate comprehensive data quality validation Python frameworks (Great Expectations) with automated checks for completeness, accuracy, consistency, uniqueness, timeliness, and validity constraints through configurable business rules.
- Develop complex data transformation models leveraging Data Build Tool (DBT) Python frameworks to ensure scale and efficient data outcomes with lineage.
- Design and implement RESTful API endpoints utilizing Swagger and OpenAPI specifications developing backend services that accept CSV, XML, JSON, and Microsoft Excel file uploads with intelligent parsing algorithms transforming diverse formats into standardized JSON representations.
- Configure ingestion schedules and monitor pipeline execution status via telemetry frameworks.
- Engineer and optimize a scalable Snowflake ecosystem by implementing sophisticated data modeling, rigorous security controls, and advanced performance tuning, while pioneering the use of advanced AI and data product capabilities using Snowflake Cortex, Snowpark, vectorized storage, and machine-learning features.
- Build and manage high-performance ETL and ELT pipelines using Delta Live Tables (DLT), leveraging its native automation for orchestration, data transformation, and quality validation.
- Manage Databricks notebooks and job configurations as code through Git integration, ensuring stable production deployments.
- Partner with data scientists, business analysts, product managers, and executive stakeholders to translate complex business requirements into technical data platform capabilities, conduct technical design reviews, architecture discussions, and knowledge transfer sessions with engineering teams.
- Provide production support for data pipeline failures performing root cause analysis and implementing preventive measures.
- Participate in agile development ceremonies.
Requirements
Bachelor’s degree in Computer Science, Engineering or a related field and two (2) years of experience in the job offered or as Data Engineer, or a related role.
Must have two (2) years of experience with:
- Python and Scala programming skills in Core Python and PySpark, including creating and supporting UDFs and modules including PyTest.
- Building and optimizing Big Data pipelines, architectures and data sets using Hadoop, Spark and HIVE.
- Data pipeline and workflow management tools including Airflow and DBT Kafka.
- Developing on Spark in a production environment including parallel execution, deciding resources and different modes of executing jobs.
- Using Hive, Yarn and Sqoop for bucketing, partitioning, tuning and handling different file formats including ORC, PARQUET and AVRO.
- Transact SQL, No-SQL and GraphQL.
- Snowflake and Pandas.
- Deployment, maintenance and administrative tasks related to Cloud, OpenStack, Docker, Kafka and Kubernetes.
- Cloud technologies including Azure and AWS.
- Databricks.
- JSON, XML and Unstructured data.
- Open API and Swagger.
Pay
$147,500 – $148,000 per year.
Benefits
- Annual discretionary bonus.
- Comprehensive healthcare.
- Leave benefits.
- Retirement benefits with a strong retirement plan.
- Tuition reimbursement.
- Support for working parents.
- Flexible Time Off (FTO) to relax, recharge and be there for the people you care about.
Schedule
BlackRock’s hybrid work model requires employees to work at least 4 days in the office per week, with the flexibility to work from home 1 day a week. Some business groups may require more time in the office due to their roles and responsibilities.