Jobs · Information Technology · Texas

Data Engineer, Forward Deployed

Applied Computing · Houston, TX · 3 days ago
HybridInformation Technology$32/hrFull-time

Data Engineering

The Data Engineer will architect and maintain pipelines that transform high-frequency time-series, lab, and historian data into a scalable Lakehouse architecture, usable for both deep learning models and real-time Large Language Models (LLMs).

  • Architect and maintain pipelines that ingest from OPC UA servers, process historians, IoT sensors, LIMS systems, alarms/events, and P&IDs.
  • Map signals to their physical processes (tags, units, hierarchies) for interpretability in AI pipelines.
  • Build and manage dual pipelines:
    • Training → large-scale historical data preparation for time-series + LLM training.
    • Inference → low-latency, real-time pipelines for anomaly detection, optimisation, and LLM search.
  • Support heterogeneous AI workloads (time-series forecasting and retrieval-augmented LLMs).
  • Tune PostgreSQL and Spark for high-throughput time-series workloads (partitioning, indexing, query optimisation).
  • Optimise pipelines for both fast analytical queries and high-efficiency model training.
  • Deploy and manage data pipelines in AWS EKS (Kubernetes) with persistent EBS-backed storage.

Technical Requirements

  • Deep expertise in PostgreSQL (partitioning, indexing, query optimisation, storage design).
  • Strong proficiency in Python for data processing, scripting, and pipeline orchestration.
  • Hands-on experience with AWS (EKS, S3, EBS, IAM, KMS, CloudWatch, etc.) for secure and scalable data pipelines.
  • Proven ability to work with Databricks and PySpark for large-scale distributed data processing.
  • Familiarity with time-series industrial data (control systems, DCS/SCADA logs, process historians).
  • Experience in unstructured data sync and management within hybrid cloud/on-prem environments.
  • Bonus: Experience working as a data engineer in oil and gas or energy environments.
  • Bonus: Knowledge of streaming frameworks (Kafka, Flink, Spark Streaming) or MLOps stacks for data versioning and lineage.

Core Responsibilities

  • Ingest & Contextualise Data
  • Build pipelines that handle real-time streaming and batch ingestion into the Lakehouse.
  • Manage synchronisation between historian archives, unstructured files, and AWS storage (S3/EBS).
  • Orchestrate Databricks Lakeflow/Connectors for integrating data into Lakebase/Lakehouse.
  • Handle secure, high-throughput transfers between historian archives and sandbox/live environments.
  • Change Tracking & Integrity
  • Detect and manage schema changes, signal drift, and inconsistencies across time.
  • Implement lineage and audit trails across Spark/Databricks and AWS pipelines.
  • Data Preparation for AI
  • Build and maintain dual pipelines for training and inference.
  • Support heterogeneous AI workloads (time-series forecasting and retrieval-augmented LLMs).
  • Database Performance & Optimisation
  • Tune PostgreSQL and Spark for high-throughput time-series workloads (partitioning, indexing, query optimisation).
  • Optimise pipelines for both fast analytical queries and high-efficiency model training.
  • Deploy and manage data pipelines in AWS EKS (Kubernetes) with persistent EBS-backed storage.

What Success Looks Like

  • Liv data streams are contextualised, queryable, and AI-ready.
  • Schema changes and signal drift are detected and handled without breaking downstream workflows.
  • Training and inference pipelines run smoothly in parallel, optimised for scale and latency.

Similar jobs

Data Engineer

Wolfe, LLCPittsburgh, PA· 4 days ago
Information Technologyapply on apply.wolfe.com

Data Engineer

Ascot GroupDallas, TX· 4 days ago
Information Technology$110k–$125k/yrapply on fa-emkq-saasfaprod1.fa.ocs.oraclecloud.com

Data Engineer

ProlaioChicago, IL· 5 days ago
Analyst$134k/yrapply on job-boards.greenhouse.io

Data Engineer

Evlo AINew York, NY· 5 days ago
RemoteEngineeringapply on evlo.ai

Data Engineer

Chenega MIOS SBUHuntsville, AL· 1 mo ago
apply on chenega.jibeapply.com

Data Engineer

HaystackWashington, DC· 1 mo ago
RemoteInformation Technology$62k–$141k/yrapply on haystack.cv