Technical Support Engineer
Work location: Chicago, USA
About This Role
We are seeking an experienced L3 Technical Support Engineer to provide onsite production support for enterprise data engineering platforms at a strategic customer location. In this role, you will serve as the primary technical interface between the customer and the broader engineering organization, owning complex incidents end-to-end and coordinating closely with offshore support and engineering teams. You will work across AWS-based data services, PySpark, AWS Glue, Databricks, Amazon Redshift, and SQL to ensure platform stability, SLA adherence, and continuous operational improvement. This is a high-visibility, customer-facing role requiring deep technical expertise, strong stakeholder management, and communication skills. The client works in the logistics domain (passenger and goods).
Responsibilities
- Provide L3 production support for data engineering applications, ETL/ELT pipelines, batch workflows, and analytical data platforms.
- Own complex incidents from initial triage through resolution, including log analysis, data validation, defect isolation, root cause analysis, recovery, and preventive-action tracking.
- Troubleshoot and resolve issues across AWS Glue, Amazon Redshift, Databricks, PySpark applications, SQL workloads, and related AWS services.
- Diagnose data mismatches, pipeline failures, performance degradation, dependency issues, access problems, and environment or configuration defects.
- Write and optimize complex SQL queries for troubleshooting, reconciliation, data-quality validation, and performance analysis.
- Debug PySpark and Spark-based workloads using execution plans, job logs, cluster metrics, Spark UI, and application diagnostics.
- Monitor production platforms proactively, identify risks, and ensure incidents and service requests are resolved within agreed SLAs.
- Lead incident bridges for critical production issues and provide timely, accurate updates to technical teams, business stakeholders, and customer leadership.
- Coordinate daily with offshore support and data engineering teams, including work allocation, technical handoffs, knowledge transfer, and status reporting.
- Partner with data engineering, platform, infrastructure, DevOps, security, QA, and vendor teams to implement permanent fixes and maintain platform stability.
- Support release, deployment, change, and post-production validation activities across development, test, and production environments.
- Prepare and maintain runbooks, troubleshooting guides, known-error records, incident reports, RCA documents, and operational dashboards.
- Identify opportunities to automate repetitive support activities, improve monitoring and alerting, and reduce manual recovery effort.
- Mentor L1/L2 support engineers and enable the offshore team to resolve recurring issues independently.
- Participate in on-call or extended-hours support rotation when required for business-critical incidents.
Requirements
Technical Skills
- Strong hands-on experience with AWS data services, particularly AWS Glue, Amazon Redshift, Amazon S3, IAM, and Amazon CloudWatch.
- Strong proficiency in PySpark, Spark SQL, and distributed data-processing concepts.
- Hands-on experience supporting and troubleshooting Databricks jobs, notebooks, workflows, clusters, and runtime environments.
- Advanced SQL skills, including complex joins, analytical functions, stored procedures, query-plan analysis, and performance tuning.
- Strong debugging and problem-solving ability across data pipelines, application code, job orchestration, databases, and cloud infrastructure.
- Experience with ETL/ELT processing, data warehousing, data lakes, data-quality controls, reconciliation, and large-volume data handling.
- Knowledge of Redshift workload management, distribution and sort keys, query optimization, connectivity, and operational troubleshooting.
- Experience using monitoring, logging, alerting, and ticketing tools for production support and incident management.
- Working knowledge of Git-based version control, CI/CD processes, deployment practices, and environment promotion.
- Familiarity with Python or shell scripting for diagnostics, automation, and operational tooling.
Professional Skills
- Excellent customer-facing communication and stakeholder-management skills.
- Proven ability to coordinate effectively between onsite stakeholders and offshore delivery teams across time zones.
- Strong ownership, prioritization, and decision-making skills in high-pressure production situations.
- Ability to explain technical issues, business impact, workarounds, and resolution plans to both technical and non-technical audiences.
- Strong documentation, analytical, and attention-to-detail skills.
- Ability to work independently at the customer site while maintaining close alignment with the wider engineering organization.
Qualifications & Experience
- 6–8 years of overall experience in technical support, production support, data engineering, or a related discipline, including substantial L3 support responsibilities.
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
- Demonstrated experience working onsite at a customer location and coordinating with offshore teams.
- Proven experience supporting business-critical data platforms in production environments.
Preferred
- AWS or Databricks certification.
- Experience with Apache Airflow, AWS MWAA, AWS Step Functions, Lambda, or similar orchestration services.
- Exposure to ServiceNow, Jira, or comparable IT service-management platforms.
- Experience defining SLAs, operational metrics, knowledge-management practices, and continuous-improvement initiatives.
- Experience in a regulated or large enterprise customer environment.