Senior Systems Engineer
Aprio · Atlanta, GA · 2 wk ago
HybridFull-time
Responsibilities
- Domain ownership: Own the operational health of one or two infrastructure domains (e.g., server platforms, virtualization, cloud infrastructure, identity, backup and recovery, specialty business systems).
- Major incident leadership: Lead major incident response: drive technical resolution, coordinate responders, communicate to stakeholders, and own the post-incident review and corrective actions.
- Patching and DR programs: Own the patching program for assigned domains — cadence, exception handling, reporting, continuous improvement. Own disaster recovery execution: maintain and exercise DR runbooks, coordinate tests, and ensure recovery objectives are met and evidenced.
- Server lifecycle at scale: Drive capacity planning, refresh cycles, configuration baselines, and decommissioning of legacy systems.
- Monitoring and observability: Operate and improve monitoring and observability across managed systems — tune alerting, eliminate noise, build dashboards, and contribute to anomaly-detection and AIOps initiatives.
- Cross-team initiatives: Lead initiatives that span Platform Engineering, Cybersecurity, Networking, and application teams — controlled rollouts, hardening efforts, platform migrations.
- Land them without breaking production. Standards and patterns: Define and document operational patterns, runbooks, and standards the team executes against and the auditors review.
- Mentorship: Pair with Systems Engineers, run technical reviews, give substantive feedback, and grow the next tier.
- Operational partnership: Be the senior partner Platform Engineering, Cybersecurity, Networking, and IT Service Management call when they need operational input. Solve problems with them, not at them.
- Specialty systems: Provide operational ownership of specialty business systems and legacy platforms supporting business-critical applications.
- Security and Audit: Apply and validate security baselines, lead remediation of high-severity findings, and keep your domains' evidence map current.
- Automation: Push toward repeatable, codified operations (IaC, automated evidence collection, scripted runbooks) instead of one-off manual work.
- On-call: Participate in and serve as senior escalation for the on-call rotation, including after-hours support for high-severity incidents, change windows, and disaster recovery events.
Qualifications
- 5+ years in systems engineering, systems administration, or infrastructure operations, including time in a senior individual contributor or technical lead capacity.
- Strong fundamentals across multiple infrastructure domains (server platforms, virtualization, cloud infrastructure, networking, identity, backup and recovery).
- Experience operating production workloads in at least one major cloud platform.
- Demonstrated experience leading major incidents and disaster recovery exercises.
- Able to produce clear architecture, operational, and decision documentation that holds up under audit and peer review.
- Excellent written and verbal communication; able to explain trade-offs across technical and business audiences in plain language.
- Comfortable mentoring less senior engineers and owning quality-of-output for one or more domains.
- Comfortable serving as senior escalation in an on-call rotation.
Preferred Qualifications
- Deep hands-on experience with Microsoft Azure (compute, networking, storage, identity, RBAC, monitoring).
- Advanced administration experience with enterprise virtualization, hybrid Active Directory, and endpoint management platforms.
- Infrastructure-as-code experience (Terraform, Bicep, ARM) and exposure to policy-as-code.
- Advanced scripting / automation (PowerShell, Python, Bash) and experience automating operational work at scale.
- Experience with observability or AIOps platforms.
- Operations or administration experience with specialty business systems (e.g., IBM iSeries / AS400).
- Experience operating within regulated environments (SOC 2, ISO 27001, HIPAA, PCI, NIST 800-171, CMMC, or similar).
- Industry certifications (Microsoft AZ-104 / AZ-305 / AZ-500, MS-102, VMware VCP, Red Hat RHCSA/RHCE, ITIL Foundation, or equivalents).
- Experience supporting a professional services, accounting, or financial services firm.
- Bachelor's degree in Computer Science, Information Systems, or related field — or equivalent applicable years of experience.