Platform Engineer I
About Blackhawk Network
Through BHN’s single global platform, businesses can tap into the world’s largest network of branded payment solutions. BHN helps businesses grow revenue, increase loyalty, motivate and reward their teams, disburse funds, and engage consumers. Branded payment solutions include the issuance and distribution of gift cards, egifts, corporate payouts and rewards, along with the technology to deliver these products in seamless, integrated ways. BHN’s network spans the globe with more than 400,000 consumer touchpoints.
Hybrid flexibility: Enjoy focused remote work plus in-person collaboration at our Coppell, TX office, providing the tools, connection, and autonomy to make a real impact.
Overview
The Operations Command Centre (OCC) is looking for a Platform Engineer with strong technical foundations, exceptional problem-solving ability, and a passion for building reliable systems. This role splits time between engineering solutions and operating production platforms—maintaining the health of BHN's production services while building automation, observability, and AI-driven capabilities to reduce incidents, improve diagnosis, and accelerate resolution.
As part of the OCC, you'll actively participate in Major Incident Management, partnering with engineering teams to diagnose and restore production services during critical incidents. Outside of incident response, you'll build dashboards, improve monitoring, develop automation, analyze operational data, and engineer intelligent tooling to continuously improve platform reliability.
This role provides exposure to large-scale distributed systems, cloud infrastructure, Kubernetes, CI/CD, observability platforms, automation, AI-assisted software development, and production engineering.
Responsibilities
Major Incident Response & Production Operations
- Participate in the 24×7 on-call rotation supporting BHN's production platforms.
- Monitor production health using modern observability platforms.
- Lead or support Major Incident bridges, coordinating technical teams during high-severity production incidents.
- Perform technical triage, identify probable causes, and drive rapid service restoration.
- Communicate clearly with engineers, leadership, and business stakeholders throughout incidents.
- Lead post-incident reviews focused on learning and continuous improvement.
- Identify recurring operational pain points and engineer permanent solutions.
Platform Engineering & Automation
- Develop automation that reduces manual operational effort.
- Build internal engineering tools that improve developer productivity and platform reliability.
- Create dashboards, alerts, health scores, and operational insights.
- Improve CI/CD pipelines and deployment safety.
- Automate operational workflows and repetitive tasks.
- Build self-service capabilities for engineering teams.
- Develop auto-remediation and self-healing capabilities.
- Continuously improve platform reliability through engineering rather than manual intervention.
Observability & Reliability Engineering
- Design alerts that detect customer-impacting issues early while minimizing alert fatigue.
- Improve platform visibility through metrics, logs, traces, and dashboards.
- Analyze production behavior to identify reliability improvements.
- Develop operational KPIs and engineering health metrics.
- Define and measure Service Level Indicators (SLIs) and Service Level Objectives (SLOs).
- Use operational data to drive engineering decisions and improve platform resilience.
AI Engineering & Intelligent Operations
- Use AI-assisted software development tools to improve engineering productivity.
- Develop AI-powered incident summarization and communication capabilities.
- Build intelligent root cause analysis and diagnostic tooling.
- Create operational copilots and engineering assistants.
- Enhance alerts with contextual intelligence.
- Automate diagnostics and operational workflows.
- Build AI-driven orchestration and auto-remediation capabilities.
- Develop engineering knowledge systems that improve troubleshooting and accelerate learning.
Experience You'll Gain
- Cloud Infrastructure (AWS)
- Kubernetes
- CI/CD Engineering
- Infrastructure as Code
- Observability & Monitoring
- Major Incident Management
- Production Operations
- Reliability Engineering (SRE)
- Automation Engineering
- AI Engineering
- Large-scale Distributed Systems
Qualifications
Required
- Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.
- Experience in Platform Engineering, DevOps, Site Reliability Engineering (SRE), Infrastructure Engineering, Technical Operations, or a similar technical role.
- Strong Linux fundamentals.
- Experience with AWS or another major cloud platform.
- Experience with Git and modern software development workflows.
- Basic scripting experience using Python, Bash, or a similar language.
- Strong analytical and troubleshooting skills, with exposure to production support or Major Incident Management.
- Excellent written and verbal communication skills.
- Strong ownership mindset with a passion for continuous improvement.
- Experience using modern AI engineering tools such as GitHub Copilot, Cursor, Claude, or similar AI-assisted development platforms.
Nice to Have
- Kubernetes
- Docker
- Jenkins
- Splunk
- New Relic
- Prometheus
- Grafana
- OpenTelemetry
- ServiceNow
- Python automation
- Infrastructure as Code (Terraform, CloudFormation, etc.)
- CI/CD engineering
- Distributed systems
- FinTech, payments, or other high-availability production environments
We seek candidates who demonstrate curiosity and adaptability in emerging technologies and have successfully implemented and utilized AI tools to enhance their work, improve processes, or deliver measurable results.
Benefits
- 401k with employer match
- Medical, dental, and vision insurance
- 12 paid holidays throughout the year
- Sick pay accrual according to state law
- Parental leave
- Life insurance
- Disability insurance
- Accident and illness insurance
- Health and dependent care flexible spending accounts
- Wellness benefits
- Paid time off for all full-time employees