Data Engineer - Web Scraping
Jobgether · United States · Today
RemoteRemoteInformation TechnologyFull-time
The Data Engineer - Web Scraping role is listed on behalf of a partner company managing all applications and next steps. This position is located in the United States.
About the role
As a Data Engineer specializing in web scraping, you will design, develop, and maintain automated data collection systems while ensuring the quality, accuracy, and availability of large-scale datasets. You will collaborate across teams to build efficient, reliable, and scalable data solutions. You will also investigate and resolve data pipeline issues and time-sensitive production incidents.
Responsibilities
- Design, develop, and maintain web scrapers for a wide range of structured and unstructured data sources.
- Clean, transform, validate, and manipulate large datasets using Python and Pandas.
- Build and maintain data ingestion pipelines into databases or data warehouses.
- Schedule, monitor, and optimize scraping workflows using orchestration tools such as Apache Airflow.
- Develop quality control checks to ensure data integrity, consistency, and availability.
- Investigate and resolve data pipeline issues and time-sensitive production incidents.
- Design and enhance internal tools, automation frameworks, and platform capabilities to improve operational efficiency.
- Work closely with cross-functional engineering teams to implement scalable and maintainable data processing workflows.
Requirements
- A Bachelor's or Master's degree in Computer Science or a related technical discipline.
- 2-4 years of professional software development experience.
- Strong programming skills in Python and solid SQL/database knowledge.
- Advanced experience using the Pandas library for data cleaning, transformation, and analysis.
- Experience working with web technologies, including HTML, JavaScript, APIs, and related protocols.
- Proven experience processing, cleaning, and transforming large datasets.
- Familiarity with web scraping frameworks and tools such as Selenium, Scrapy, XPath, Fiddler, or Postman.
- Experience with workflow orchestration tools such as Apache Airflow or similar platforms.
- Knowledge of Docker containerization; Kubernetes experience is an advantage.
- Experience working with cloud services, particularly AWS technologies such as S3, RDS, Lambda, SNS, or SQS, is preferred.
- Strong analytical thinking, attention to detail, communication skills, and a passion for automation and continuous improvement.
- Opportunity to work on challenging projects supporting a leading global asset management environment.
- High level of ownership and autonomy in a collaborative, team-oriented culture.
- Exposure to modern data engineering, web scraping, cloud, and automation technologies.
- Collaborative environment with experienced engineering, product, and data professionals.
- Opportunities for continuous learning, professional growth, and technical skill development.
- Merit-driven culture that values innovation, initiative, and individual contributions.
- Flexible, technology-focused environment with opportunities to work on impactful data products.