Infrastructure Reliability Engineer
Pinnacle Arc LLC · Middlesex County, NJ · 6 days ago
EngineeringFull-time
About the role
We are looking for a hands-on Infrastructure Reliability Engineer to improve reliability, automation, security, and operability across enterprise infrastructure—spanning cloud and on-prem/hybrid environments. This role focuses on Terraform + Ansible automation, Windows and Unix/Linux operations, observability, incident/postmortems, and strong vulnerability remediation ownership across servers, platforms, and containerized workloads.
Key Responsibilities
- Security & Vulnerability Remediation
- Lead remediation of infrastructure vulnerabilities across Windows, Linux, middleware, and supporting Frontier AI platform components.
- Drive closure of control findings, audit items, and cyber remediation commitments within agreed timelines.
- Partner with Cybersecurity, Infrastructure, and Application Development teams to identify, prioritize, and remediate vulnerabilities at scale.
- Establish sustainable patching, upgrade, and lifecycle management processes to reduce recurring findings.
- Infrastructure Automation (Terraform + Ansible)
- Build and maintain Infrastructure as Code using Terraform across cloud and/or virtualized environments.
- Automate configuration, provisioning, patching, and deployments using Ansible across Linux/Unix and Windows estates.
- Standardize environments (dev/test/stage/prod), build reusable modules/playbooks, and enforce configuration consistency (prevent drift).
- Hybrid / Enterprise Infrastructure Operations
- Operate and troubleshoot infrastructure components end-to-end including compute (VMs/servers), virtualization (VMware or equivalent), networking (DNS, routing, VPNs, proxies), load balancers, and security controls.
- Partner with application teams to ensure infrastructure supports scalable, reliable application delivery.
- CI/CD, Observability & Reliability (SRE)
- Implement and support CI/CD automation (Jenkins, Spinnaker, Cloud Deployment) and enable safe release patterns (blue/green, canary, automated rollback).
- Build and maintain monitoring/alerting dashboards using Prometheus, Grafana, Dynatrace, and Splunk.
- Improve MTTR through better telemetry and define/drive reliability outcomes using SLIs/SLOs.
Must-Have Skills
- Strong experience in SRE / DevOps / Infrastructure Engineering / Production Operations.
- Proven hands-on automation with Terraform and Ansible in real production environments.
- Strong administration and troubleshooting across Windows and Unix/Linux estates.
- Experience supporting enterprise infrastructure areas (compute, network, storage, load balancing, security controls).
- Demonstrated experience with vulnerability remediation (scanning, triage, patching, and verification).
- Hands-on with vulnerability management tools (Qualys, Tenable, Rapid7, etc.).
- Audit and controls remediation experience across Cyber, Risk, Controls, Infrastructure, and Application teams.
- Observability tooling: Prometheus, Grafana, Dynatrace, Splunk, and cloud-native monitoring.
Required Qualifications
- 13+ years of professional experience in SRE, DevOps, Infrastructure Engineering, or Production Operations.
- Proven hands-on automation with Terraform and Ansible in real production environments.
- Strong administration and troubleshooting across Windows and Unix/Linux.
- Experience supporting enterprise infrastructure areas (compute, network, storage, load balancing, security controls).
- Practical incident management experience (on-call, RCA/postmortems, operational improvements).
- Demonstrated experience with vulnerability remediation (actual patching and verification).
Preferred Qualifications
- Prior EX-JPMC experience required.
- Audit and controls remediation experience.
- Experience with containerized workloads and cloud-native infrastructure.
Location: NJ / OH (Onsite)
Work Authorization: OPT, USC, H1B, TN, H4EAD, L2EAD (Strictly No GC, GCEAD and CPT)
Employment Type: C2C