Data Engineer
Bespoke Technologies, Inc. · Chantilly, VA · 1 wk ago
Information TechnologyFull-time
Must have an active TS/SCI clearance to apply. Candidates without an active security clearance will not be considered.
About the Role
Bespoke Technologies is seeking a Software Developer to provide ETL, Data Engineering, and Full-Stack Software Development support.
Required Skills
- Experience designing and maintaining enterprise-grade ETL/ELT pipelines, both batch and real-time.
- Front-end development and implementation skills using React, Next.js, or similar.
- Back-end development using Python, Java, Scala, and microservices architecture.
- Experience with API design.
- Experience with containerization using Docker.
- Experience with CI/CD pipelines.
- Experience with infrastructure-as-code patterns.
- Experience with probabilistic models, risk scoring, Bayesian inference, Monte Carlo simulation, and probabilistic graphical models.
- Experience applying statistical modeling tools.
- Experience designing cloud-native architectures using cloud services such as AWS, Google, IBM, and Oracle.
- Experience designing and operating big data systems.
- Experience building and optimizing performance of large-scale graph databases (tens of billions of edges) using DynamoDB or new enhanced capabilities.
- Experience developing and operating graph traversal capabilities using data graphing tool traversal capabilities built upon Apache Gremlin or new enhanced capabilities.
- Experience developing and operating NoSQL solutions to complex big data applications.
- Experience in data modeling for performance, partition sharding, record/event aggregation workflows, stream processing, and metrics gathering.
- Experience designing and operating large-scale serverless geospatial indexes built with GeoMESA.
- Experience with partition and sort key design and implementation to ensure consistent performance.
- Experience with aggregation operations to de-duplicate records on continuous data feeds.
- Subject matter expertise with relational databases to NoSQL.
- Experience building and operating high-performance data processing pipelines using Lambda, Step Functions, and PySpark.
- Experience building high-quality User Interface/User experiences with the React framework and WebGL.
- Experience designing and operating large-scale graph databases using Apache Cassandra.
- Experience performing in-depth technical analysis of large-scale graph databases to develop implementation strategies for search optimizations.
- Experience developing technical capabilities for processing, persistence, and search of datasets that are collected or maintained using standards common in the Sponsor's community.
- Experience facilitating engineering discussions across teams representing multiple stakeholders to develop and execute implementation strategies that meet mission needs.
- Experience developing Machine Learning Operations (MLOps) pipelines for large-scale applications.
- Experience maintaining configuration of software using configuration management resources such as GitHub.
- Experience designing, building, and operating big data systems, such as persistence, partitioning, indexing, at scale of trillions of records/events.
- Experience with Niagara Files (NiFi) applications or new enhanced capabilities.
- Experience developing and operating Kubernetes infrastructure.
- Experience supporting engineering efforts that contribute to delivery of capabilities such as datasets and functionality like communications, geospatial workflows.
- Experience implementing DevSecOps and agile development in production environments.
- Experience with agile software development and testing.
- Experience with federal security, regulatory, and compliance requirements and security accreditation package development.
- Experience with data security and governance using centralized security controls like LDAP, encrypting the data, and auditing access to the data.
- Experience with specialized technologies optimized for particular data use, such as relational databases, NoSQL databases (Cassandra), or object storage.
- Experience with Apache, TINKERPOP, GREMLIN, and/or JANUSGRAPH to design, develop, implement, and maintain systems.
- Knowledge of Graph Database to design, develop, implement, and maintain systems.
- Experience with C or C++ to write interfaces.
- Experience using centralized security controls like LDAP, encrypting data, and auditing access to data.
- Databases: Postgres, MariaDB, ELK, Minio, AWS S3, Neo4j, MongoDB, NoSQL.
- Languages: Python (PyPI libraries).
- Operating Systems: CentOS 7, Rocky Linux 8.
- Orchestration: Kubernetes, Docker, Docker-Compose, Docker-Swarm.
- Development Tools: VSCode, GitLab, JupyterHub/Notebooks, MATLAB.
- Environments: Large collaboration and development environments.
- Data types: Unstructured, structured, or semi-structured data, including CSV, JSON, JSONL, AVRO, Protocol Buffers, Parquet, etc.
Desired Skills
- Experience designing cloud-native architectures using Sponsor's cloud services.
- Experience designing and operating big data systems within the Sponsor's policy and regulatory environment.
- Experience developing and operating graph traversal capabilities using the Sponsor's data graphing tool traversal capabilities built upon Apache Gremlin.
- Experience building and operating high-performance data processing pipelines using Lambda, Step Functions, and PySpark on the Sponsor's infrastructure with EMR.
- Experience working with the Sponsor's enterprise services used for Data Management, including the enterprise catalog service (and associated APIs) and Policy Decision Points (PDPs).
- Experience developing Machine Learning Operations (MLOps) pipelines for large-scale applications in the Sponsor's environment.
- Understanding of IT Service Management and common SLA measurements.
- Experience presenting solutions, requirements, and presentations to diverse audiences.
- Experience working with container orchestration technologies such as AWS ECS, AWS Fargate, and Kubernetes or other enhanced capabilities available.
- Experience managing large operational cloud environments spanning multiple tenants using Multi-Account management, AWS Well-Architected Best Practices, and AWS Organization Units/Service Control Policies (OU/SCP).
- Experience with microservices such as building decoupled systems, utilizing RESTful endpoints, and lightweight systems.
- Experience in total systems perspectives, including a technical understanding of systems and applications relationships, dependencies, and requirements of hardware and software components.
- Experience consulting with customers to determine present and future user needs.
- Experience providing frequent contact with customers, traceability within program documents, and the overall computing environment and architecture.
Desired Certifications
- AWS Certified Solutions Architect
- AWS Machine Learning Certification(s)
- Agile certification
- Azure
- Security+
- GSEC
- CCNA