Jobs · Engineering · New Jersey

Director of Software Engineering - Payment Data Platform Infrastructure

hackajob · Jersey City, NJ · Yesterday
On-siteEngineeringFull-time

Job Responsibilities

  • Provides overall direction, oversight, and coaching for a team of engineering managers and senior technologists spanning infrastructure engineering, SRE, and SecOps for both BRIE and NEO.
  • Owns the reliability strategy for the platform estate - SLOs, error budgets, capacity planning, and disaster recovery.
  • Sustains active-active, multi-region operation across AWS and on-prem while meeting the platform's high-availability commitments.
  • Owns the security operations agenda: threat detection and response, vulnerability and patch management, secrets and key management, and evidence for audit, risk, and regulatory reviews, maintaining CPOF compliance across all environments.
  • Champions infrastructure-as-code, immutable deployments, and platform automation so provisioning, scaling, and remediation are repeatable, reviewable, and auditable.
  • Sets direction and governance for agentic AI-enabled engineering and SDLC/TLM automation within a technical area to drive measurable improvements in speed, quality, and operational outcomes (e.g., AI-orchestrated delivery workflows, release readiness controls, automated test modernization, and incident triage acceleration), while establishing guardrails for validation, security, resiliency, traceability, and reuse across teams.
  • Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation and support capacity unlock initiatives at scale.
  • Led incident command for major events, drives blameless postmortems, and closes the loop on systemic remediation to raise overall operational stability.
  • Makes decisions that influence resourcing, budget, tooling, and vendor selection across the infrastructure organization, and is accountable for those outcomes.
  • Promotes evaluation sessions with cloud providers, vendors, startups, and internal teams to probe architectural designs, technical credentials, and applicability within existing systems and information architecture.
  • Pairs with platform, product, data, and InfoSec stakeholders, and communicates reliability, security, and cost trade-offs to senior leadership.

Required Qualifications, Capabilities, And Skills

  • Formal training or certification on software engineering concepts and 8+ years applied experience, including significant time leading infrastructure, SRE, or platform organizations, with experience managing managers.
  • Experience leading teams of technologists and managing budget, resourcing, and delivery across multiple concurrent workstreams.
  • Hands-on background in large-scale distributed systems: system design, application development, testing, and operational stability.
  • Deep expertise operating production systems across public cloud and on-premises data centers, including multi-region, active-active resilience and disaster recovery.
  • Demonstrated ownership of a security posture in a regulated environment - SecOps, identity and access, secrets management, audit, and regulatory compliance.
  • Experience leading adoption of agentic AI-enabled engineering practices (using enterprise-authorized tools within the work environment) across teams, including defining operating expectations (human-in-the-loop validation, quality gates), measuring outcomes, and ensuring secure handling of sensitive inputs/outputs.
  • Strong understanding of responsible AI use and control expectations in engineering workflows, including data sensitivity, resiliency/security implications, and governance; ability to influence leaders on safe scaling patterns and reuse.
  • Proficient in all aspects of the Software Development Life Cycle.
  • Advanced understanding of agile methodologies such as CI/CD, Application Resiliency, and Security.
  • Advanced in one or more programming language(s), with the ability to guide critical technical decisions.
  • Demonstrated proficiency in software applications and technical processes within a technical discipline (e.g., cloud, artificial intelligence, machine learning, distributed data systems).

Preferred Qualifications, Capabilities, And Skills

  • Deep AWS experience across multi-account, multi-region architectures, paired with on-premises data center operations.
  • Expert-level Kubernetes and containerized platform operations, including multi-tenant isolation (siloed and pooled) and noisy-neighbor controls.
  • Proficiency with Infrastructure as Code (Terraform) and GitOps-driven, immutable deployment pipelines.
  • Experience operating policy and authorization at the infrastructure layer (e.g., OPA/Rego, OpenFGA) and policy-as-code workflows and experience operating Databricks (Spark) and real-time stream processing with Apache Flink at production scale.
  • Strong observability and SRE tooling background (e.g., Splunk, OpenTelemetry, Prometheus/Grafana) with SLO/error-budget practice.
  • Experience running GPU compute fleets for ML serving and training (e.g., NVIDIA L40S/H100, SageMaker) and associated capacity and cost governance and experience with regulated-environment compliance frameworks and evidencing controls to risk, audit, and InfoSec.
  • Familiarity building and running Java Spring Boot GraphQL services for high-availability, low-latency APIs and familiarity with the data platform stack - lakehouse and open table formats (Apache Iceberg), streaming and batch pipelines, and OLAP/columnar stores.

Similar jobs