Data Engineer, Forward Deployed
Applied Computing · Houston, TX · 3 days ago
HybridInformation Technology$32/hrFull-time
Data Engineering
The Data Engineer will architect and maintain pipelines that transform high-frequency time-series, lab, and historian data into a scalable Lakehouse architecture, usable for both deep learning models and real-time Large Language Models (LLMs).
- Architect and maintain pipelines that ingest from OPC UA servers, process historians, IoT sensors, LIMS systems, alarms/events, and P&IDs.
- Map signals to their physical processes (tags, units, hierarchies) for interpretability in AI pipelines.
- Build and manage dual pipelines:
- Training → large-scale historical data preparation for time-series + LLM training.
- Inference → low-latency, real-time pipelines for anomaly detection, optimisation, and LLM search.
- Support heterogeneous AI workloads (time-series forecasting and retrieval-augmented LLMs).
- Tune PostgreSQL and Spark for high-throughput time-series workloads (partitioning, indexing, query optimisation).
- Optimise pipelines for both fast analytical queries and high-efficiency model training.
- Deploy and manage data pipelines in AWS EKS (Kubernetes) with persistent EBS-backed storage.
Technical Requirements
- Deep expertise in PostgreSQL (partitioning, indexing, query optimisation, storage design).
- Strong proficiency in Python for data processing, scripting, and pipeline orchestration.
- Hands-on experience with AWS (EKS, S3, EBS, IAM, KMS, CloudWatch, etc.) for secure and scalable data pipelines.
- Proven ability to work with Databricks and PySpark for large-scale distributed data processing.
- Familiarity with time-series industrial data (control systems, DCS/SCADA logs, process historians).
- Experience in unstructured data sync and management within hybrid cloud/on-prem environments.
- Bonus: Experience working as a data engineer in oil and gas or energy environments.
- Bonus: Knowledge of streaming frameworks (Kafka, Flink, Spark Streaming) or MLOps stacks for data versioning and lineage.
Core Responsibilities
- Ingest & Contextualise Data
- Build pipelines that handle real-time streaming and batch ingestion into the Lakehouse.
- Manage synchronisation between historian archives, unstructured files, and AWS storage (S3/EBS).
- Orchestrate Databricks Lakeflow/Connectors for integrating data into Lakebase/Lakehouse.
- Handle secure, high-throughput transfers between historian archives and sandbox/live environments.
- Change Tracking & Integrity
- Detect and manage schema changes, signal drift, and inconsistencies across time.
- Implement lineage and audit trails across Spark/Databricks and AWS pipelines.
- Data Preparation for AI
- Build and maintain dual pipelines for training and inference.
- Support heterogeneous AI workloads (time-series forecasting and retrieval-augmented LLMs).
- Database Performance & Optimisation
- Tune PostgreSQL and Spark for high-throughput time-series workloads (partitioning, indexing, query optimisation).
- Optimise pipelines for both fast analytical queries and high-efficiency model training.
- Deploy and manage data pipelines in AWS EKS (Kubernetes) with persistent EBS-backed storage.
What Success Looks Like
- Liv data streams are contextualised, queryable, and AI-ready.
- Schema changes and signal drift are detected and handled without breaking downstream workflows.
- Training and inference pipelines run smoothly in parallel, optimised for scale and latency.