Director of Software Engineering - Payment Data Platform Infrastructure
hackajob · Jersey City, NJ · Yesterday
On-siteEngineeringFull-time
Job Responsibilities
- Provides overall direction, oversight, and coaching for a team of engineering managers and senior technologists spanning infrastructure engineering, SRE, and SecOps for both BRIE and NEO.
- Owns the reliability strategy for the platform estate - SLOs, error budgets, capacity planning, and disaster recovery.
- Sustains active-active, multi-region operation across AWS and on-prem while meeting the platform's high-availability commitments.
- Owns the security operations agenda: threat detection and response, vulnerability and patch management, secrets and key management, and evidence for audit, risk, and regulatory reviews, maintaining CPOF compliance across all environments.
- Champions infrastructure-as-code, immutable deployments, and platform automation so provisioning, scaling, and remediation are repeatable, reviewable, and auditable.
- Sets direction and governance for agentic AI-enabled engineering and SDLC/TLM automation within a technical area to drive measurable improvements in speed, quality, and operational outcomes (e.g., AI-orchestrated delivery workflows, release readiness controls, automated test modernization, and incident triage acceleration), while establishing guardrails for validation, security, resiliency, traceability, and reuse across teams.
- Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation and support capacity unlock initiatives at scale.
- Led incident command for major events, drives blameless postmortems, and closes the loop on systemic remediation to raise overall operational stability.
- Makes decisions that influence resourcing, budget, tooling, and vendor selection across the infrastructure organization, and is accountable for those outcomes.
- Promotes evaluation sessions with cloud providers, vendors, startups, and internal teams to probe architectural designs, technical credentials, and applicability within existing systems and information architecture.
- Pairs with platform, product, data, and InfoSec stakeholders, and communicates reliability, security, and cost trade-offs to senior leadership.
Required Qualifications, Capabilities, And Skills
- Formal training or certification on software engineering concepts and 8+ years applied experience, including significant time leading infrastructure, SRE, or platform organizations, with experience managing managers.
- Experience leading teams of technologists and managing budget, resourcing, and delivery across multiple concurrent workstreams.
- Hands-on background in large-scale distributed systems: system design, application development, testing, and operational stability.
- Deep expertise operating production systems across public cloud and on-premises data centers, including multi-region, active-active resilience and disaster recovery.
- Demonstrated ownership of a security posture in a regulated environment - SecOps, identity and access, secrets management, audit, and regulatory compliance.
- Experience leading adoption of agentic AI-enabled engineering practices (using enterprise-authorized tools within the work environment) across teams, including defining operating expectations (human-in-the-loop validation, quality gates), measuring outcomes, and ensuring secure handling of sensitive inputs/outputs.
- Strong understanding of responsible AI use and control expectations in engineering workflows, including data sensitivity, resiliency/security implications, and governance; ability to influence leaders on safe scaling patterns and reuse.
- Proficient in all aspects of the Software Development Life Cycle.
- Advanced understanding of agile methodologies such as CI/CD, Application Resiliency, and Security.
- Advanced in one or more programming language(s), with the ability to guide critical technical decisions.
- Demonstrated proficiency in software applications and technical processes within a technical discipline (e.g., cloud, artificial intelligence, machine learning, distributed data systems).
Preferred Qualifications, Capabilities, And Skills
- Deep AWS experience across multi-account, multi-region architectures, paired with on-premises data center operations.
- Expert-level Kubernetes and containerized platform operations, including multi-tenant isolation (siloed and pooled) and noisy-neighbor controls.
- Proficiency with Infrastructure as Code (Terraform) and GitOps-driven, immutable deployment pipelines.
- Experience operating policy and authorization at the infrastructure layer (e.g., OPA/Rego, OpenFGA) and policy-as-code workflows and experience operating Databricks (Spark) and real-time stream processing with Apache Flink at production scale.
- Strong observability and SRE tooling background (e.g., Splunk, OpenTelemetry, Prometheus/Grafana) with SLO/error-budget practice.
- Experience running GPU compute fleets for ML serving and training (e.g., NVIDIA L40S/H100, SageMaker) and associated capacity and cost governance and experience with regulated-environment compliance frameworks and evidencing controls to risk, audit, and InfoSec.
- Familiarity building and running Java Spring Boot GraphQL services for high-availability, low-latency APIs and familiarity with the data platform stack - lakehouse and open table formats (Apache Iceberg), streaming and batch pipelines, and OLAP/columnar stores.