Hadoop PySpark Data Engineer
About the role
Infosys Data and Analytics (DNA) is at the forefront of transforming data into actionable insights, driving business growth and operational efficiency. We specialize in leveraging advanced AI and analytics to create innovative solutions that address complex business challenges. Our team is dedicated to pioneering the future of data-driven decision-making, enabling organizations to unlock new opportunities and achieve sustainable success.
As a Technology Consultant 2, you will work with cutting-edge technologies, collaborate with industry experts, and contribute to transformative projects that shape the future of business. We foster a culture of continuous learning and growth, ensuring team members thrive in a dynamic and supportive environment.
Responsibilities
- Contribute to the requirements elicitation process by documenting assigned parts of business requirements, in line with guidance provided.
- Facilitate software application design discussions and document design decisions to guide the technical team toward building software solutions.
- Participate in coding and integrate new features or updates into existing applications, focusing on maintaining system stability.
- Conduct code reviews, make changes to the codebase, and maintain code repositories.
- Implement test strategies, analyze results, and coordinate bug fixes to uphold software quality standards.
- Develop user training programs, documentation, and support frameworks to ensure a smooth transition to new software applications.
- Actively participate in resolving production issues and recommend preventive strategies to enhance system reliability.
- Maintain detailed records of code, testing techniques, and support activities to enrich the knowledge base and assist similar projects.
Requirements
- Experience in data warehousing architectural approaches.
- Sound understanding and experience with the Hadoop ecosystem (Cloudera).
- Experience working with a Big Data implementation in a production environment.
- Experience with Big Data technologies like Hadoop, Hive, Spark, Python, etc.
- Experience in Python and Unix shell scripting.
- Experience with orchestration tools like Autosys or Airflow.
- Sound knowledge of relational databases (SQL) and experience with large SQL-based systems.
- Understanding of Agile methodologies and technologies.
Qualifications
- Bachelor’s degree or foreign equivalent required from an accredited institution. Three years of progressive experience in the specialty may substitute for each year of education.
- This position may require relocation and/or travel to the work/project location.
- All applicants authorized to work in the United States are encouraged to apply.
Preferred Skills
- Ability to understand and explore constantly evolving tools within the Hadoop ecosystem and apply them appropriately to relevant problems.
- Knowledge and experience with Cloud and containerization technologies.
- Experience with data visualization tools like Tableau.
Benefits
- Medical, Dental, Vision, and Life Insurance.
- Long-term and Short-term Disability.
- Health and Dependent Care Reimbursement Accounts.
- Accident, Critical Illness, Hospital Indemnity, and Legal Insurance.
- 401(k) plan with contributions dependent on salary level.
- Paid holidays plus Paid Time Off.
Pay
Estimated annual compensation range for candidates based in Jersey City, NJ: $82,419 to $126,600.