Big Data Application Developer
Bright Vision Technologies · Murray Hill, NY · 2 wk ago
Engineering$100k–$150k/yrFull-time
About the role
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Responsibilities
- Design, develop, and operate end-to-end big-data pipelines on Hadoop, ingesting data from a diverse mix of relational, file-based, streaming, and API-driven sources.
- Build robust ETL/ELT workflows using Apache Spark, Hive, Pig, and Sqoop, with strong attention to data quality, idempotency, error handling, and recoverability.
- Develop high-throughput streaming data pipelines using Kafka, Spark Streaming, or Flink, and integrate them with downstream analytical and operational systems.
- Optimize Spark and MapReduce jobs through careful tuning of partitioning, memory, serialization, and skew handling to meet demanding SLAs at minimal cost.
- Design and maintain data models and storage layouts on HDFS, Hive, HBase, and modern lakehouse formats (Parquet, ORC, Delta, Iceberg, Hudi) to balance flexibility and performance.
- Implement data governance, lineage, and quality controls in collaboration with data governance and security teams.
- Build robust monitoring, alerting, and logging strategies for big-data pipelines, including job-level SLAs and proactive failure detection.
- Partner with data scientists and analysts to deliver curated, reliable, and well-documented datasets that accelerate their work.
- Automate pipeline orchestration using Airflow, Oozie, or similar workflow engines, with clean dependency management and clear ownership boundaries.
- Continuously evaluate and adopt new technologies in the big-data and cloud ecosystem (EMR, Databricks, Snowflake, BigQuery) where they offer meaningful improvements.
- Lead performance reviews and architecture audits of existing pipelines, proposing concrete refactoring and optimization initiatives.
- Document data architectures, schemas, pipeline behaviors, and operational runbooks in a way that makes the platform supportable as the team scales.
- Mentor junior engineers and contribute to the team’s engineering standards and best practices.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related technical discipline.
- Five or more years of professional experience designing and operating big-data pipelines on Hadoop.
- Strong hands-on expertise with Apache Spark (Scala, Python, or Java) in production environments.
- Solid experience with Hive, HDFS, Sqoop, HBase, and the broader Hadoop ecosystem.
- Hands-on experience with streaming data platforms such as Kafka, Spark Streaming, or Flink.
- Strong SQL skills and experience working with both relational and NoSQL data stores.
- Experience with workflow orchestration tools such as Airflow or Oozie.
- Solid understanding of distributed systems concepts, including partitioning, replication, and fault tolerance.
- Strong scripting skills in Python or Shell.
- Excellent troubleshooting, debugging, and documentation skills.
Qualifications
- Required: Bachelor’s degree in Computer Science, Engineering, or a related technical discipline.
- Required: Five or more years of professional experience designing and operating big-data pipelines on Hadoop.
- Required: Strong hands-on expertise with Apache Spark (Scala, Python, or Java) in production environments.
- Required: Solid experience with Hive, HDFS, Sqoop, HBase, and the broader Hadoop ecosystem.
- Required: Hands-on experience with streaming data platforms such as Kafka, Spark Streaming, or Flink.
- Required: Strong SQL skills and experience working with both relational and NoSQL data stores.
- Required: Experience with workflow orchestration tools such as Airflow or Oozie.
- Required: Solid understanding of distributed systems concepts, including partitioning, replication, and fault tolerance.
- Required: Strong scripting skills in Python or Shell.
- Required: Excellent troubleshooting, debugging, and documentation skills.
Skills
- Apache Spark (Scala, Python, or Java)
- Hive, HDFS, Sqoop, HBase
- Kafka, Spark Streaming, or Flink
- SQL
- Workflow orchestration tools (Airflow, Oozie)
- Distributed systems concepts (partitioning, replication, fault tolerance)
- Python or Shell scripting
Benefits
Not specified
Pay
$100,000–$150,000 Annually
Schedule
Not specified
Benefits
Not specified
Contact
To apply, please send your resume to boon@bvteck.com or contact us at (908) 650-6699. Learn more about Bright Vision Technologies at www.bvteck.com.
Bright Vision Technologies is an Equal Opportunity Employer.