Jobs · Information Technology · Virginia

Data Engineer

Bespoke Technologies, Inc. · Chantilly, VA · 1 wk ago
Information TechnologyFull-time

Must have an active TS/SCI clearance to apply. Candidates without an active security clearance will not be considered.

About the Role

Bespoke Technologies is seeking a Software Developer to provide ETL, Data Engineering, and Full-Stack Software Development support.

Required Skills

  • Experience designing and maintaining enterprise-grade ETL/ELT pipelines, both batch and real-time.
  • Front-end development and implementation skills using React, Next.js, or similar.
  • Back-end development using Python, Java, Scala, and microservices architecture.
  • Experience with API design.
  • Experience with containerization using Docker.
  • Experience with CI/CD pipelines.
  • Experience with infrastructure-as-code patterns.
  • Experience with probabilistic models, risk scoring, Bayesian inference, Monte Carlo simulation, and probabilistic graphical models.
  • Experience applying statistical modeling tools.
  • Experience designing cloud-native architectures using cloud services such as AWS, Google, IBM, and Oracle.
  • Experience designing and operating big data systems.
  • Experience building and optimizing performance of large-scale graph databases (tens of billions of edges) using DynamoDB or new enhanced capabilities.
  • Experience developing and operating graph traversal capabilities using data graphing tool traversal capabilities built upon Apache Gremlin or new enhanced capabilities.
  • Experience developing and operating NoSQL solutions to complex big data applications.
  • Experience in data modeling for performance, partition sharding, record/event aggregation workflows, stream processing, and metrics gathering.
  • Experience designing and operating large-scale serverless geospatial indexes built with GeoMESA.
  • Experience with partition and sort key design and implementation to ensure consistent performance.
  • Experience with aggregation operations to de-duplicate records on continuous data feeds.
  • Subject matter expertise with relational databases to NoSQL.
  • Experience building and operating high-performance data processing pipelines using Lambda, Step Functions, and PySpark.
  • Experience building high-quality User Interface/User experiences with the React framework and WebGL.
  • Experience designing and operating large-scale graph databases using Apache Cassandra.
  • Experience performing in-depth technical analysis of large-scale graph databases to develop implementation strategies for search optimizations.
  • Experience developing technical capabilities for processing, persistence, and search of datasets that are collected or maintained using standards common in the Sponsor's community.
  • Experience facilitating engineering discussions across teams representing multiple stakeholders to develop and execute implementation strategies that meet mission needs.
  • Experience developing Machine Learning Operations (MLOps) pipelines for large-scale applications.
  • Experience maintaining configuration of software using configuration management resources such as GitHub.
  • Experience designing, building, and operating big data systems, such as persistence, partitioning, indexing, at scale of trillions of records/events.
  • Experience with Niagara Files (NiFi) applications or new enhanced capabilities.
  • Experience developing and operating Kubernetes infrastructure.
  • Experience supporting engineering efforts that contribute to delivery of capabilities such as datasets and functionality like communications, geospatial workflows.
  • Experience implementing DevSecOps and agile development in production environments.
  • Experience with agile software development and testing.
  • Experience with federal security, regulatory, and compliance requirements and security accreditation package development.
  • Experience with data security and governance using centralized security controls like LDAP, encrypting the data, and auditing access to the data.
  • Experience with specialized technologies optimized for particular data use, such as relational databases, NoSQL databases (Cassandra), or object storage.
  • Experience with Apache, TINKERPOP, GREMLIN, and/or JANUSGRAPH to design, develop, implement, and maintain systems.
  • Knowledge of Graph Database to design, develop, implement, and maintain systems.
  • Experience with C or C++ to write interfaces.
  • Experience using centralized security controls like LDAP, encrypting data, and auditing access to data.
  • Databases: Postgres, MariaDB, ELK, Minio, AWS S3, Neo4j, MongoDB, NoSQL.
  • Languages: Python (PyPI libraries).
  • Operating Systems: CentOS 7, Rocky Linux 8.
  • Orchestration: Kubernetes, Docker, Docker-Compose, Docker-Swarm.
  • Development Tools: VSCode, GitLab, JupyterHub/Notebooks, MATLAB.
  • Environments: Large collaboration and development environments.
  • Data types: Unstructured, structured, or semi-structured data, including CSV, JSON, JSONL, AVRO, Protocol Buffers, Parquet, etc.

Desired Skills

  • Experience designing cloud-native architectures using Sponsor's cloud services.
  • Experience designing and operating big data systems within the Sponsor's policy and regulatory environment.
  • Experience developing and operating graph traversal capabilities using the Sponsor's data graphing tool traversal capabilities built upon Apache Gremlin.
  • Experience building and operating high-performance data processing pipelines using Lambda, Step Functions, and PySpark on the Sponsor's infrastructure with EMR.
  • Experience working with the Sponsor's enterprise services used for Data Management, including the enterprise catalog service (and associated APIs) and Policy Decision Points (PDPs).
  • Experience developing Machine Learning Operations (MLOps) pipelines for large-scale applications in the Sponsor's environment.
  • Understanding of IT Service Management and common SLA measurements.
  • Experience presenting solutions, requirements, and presentations to diverse audiences.
  • Experience working with container orchestration technologies such as AWS ECS, AWS Fargate, and Kubernetes or other enhanced capabilities available.
  • Experience managing large operational cloud environments spanning multiple tenants using Multi-Account management, AWS Well-Architected Best Practices, and AWS Organization Units/Service Control Policies (OU/SCP).
  • Experience with microservices such as building decoupled systems, utilizing RESTful endpoints, and lightweight systems.
  • Experience in total systems perspectives, including a technical understanding of systems and applications relationships, dependencies, and requirements of hardware and software components.
  • Experience consulting with customers to determine present and future user needs.
  • Experience providing frequent contact with customers, traceability within program documents, and the overall computing environment and architecture.

Desired Certifications

  • AWS Certified Solutions Architect
  • AWS Machine Learning Certification(s)
  • Agile certification
  • Azure
  • Security+
  • GSEC
  • CCNA

Similar jobs

Data Engineer

IBMYorktown Heights, NY· 2 wk ago
Information Technologyapply on ibmglobal.avature.net

Data Engineer

DLA PiperHouston, TX· 2 wk ago
Information Technology$101k–$160k/yrapply on dlapiper.wd1.myworkdayjobs.com

Data Engineer

CodeVertex TalentOhio, United States· 1 wk ago
RemoteAnalyst$125k–$180k/yrapply on tally.so

Data Engineer

CodeVertex SystemNorth Carolina, United States· 1 wk ago
RemoteAnalyst$125k–$180k/yrapply on tally.so

Data Engineer

HaystackCalifornia, United States· 2 wk ago
Information Technology$99k–$225k/yrapply on haystack.cv

Data Engineer

Security BenefitTopeka, KS· 1 wk ago
$60/hrapply on jobs.dayforcehcm.com