Senior Systems Engineer - Delivery DevOps
Intuit · Atlanta, GA · 1 mo ago
On-siteInformation TechnologyFull-time
Responsibilities
- Systems and Infrastructure Engineering
- Design, build, and operate the AWS cloud infrastructure the platform runs on, managing the sending fleet across AWS environments, including AMI and image build pipelines, instance and capacity management, autoscaling, and running the fleet cost-effectively.
- Develop and manage Infrastructure-as-Code (for example, Terraform, Ansible, or Puppet) to provision, configure, and maintain scalable, secure server and network environments.
- Design and manage system architecture across the sending path, including networking, load balancing, and security configuration, to keep production applications and services healthy.
- Develop and maintain CI/CD pipelines and release automation behind the platform (Jenkins, Artifactory and RPM packaging, containerized services, and Kubernetes/ArgoCD-based delivery), automating deployment workflows so releases to production are reliable, repeatable, and secure.
- Configure and maintain containerized environments (Docker, Kubernetes) for application deployment, scaling, and orchestration.
- Configure and tune MTA software and host-level parameters to manage throughput, queueing, and delivery behavior across the fleet.
Operational Excellence and Reliability
- Take part in an on-call rotation, available to handle Delivery-related operational issues outside normal business hours, to keep the platform up and the business running.
- Monitor production systems and infrastructure for performance, availability, and security, and own the metrics, dashboards, and alerting that give the team and stakeholders visibility into fleet and sending health (for example, BigQuery, OpenSearch, and Splunk).
- Respond to fleet, host, and delivery incidents, diagnosing issues that span the MTA software, the hosts, the AWS environment, and the network.
- Write scripts and automation tooling to streamline infrastructure provisioning, system configuration, and operational processes, and to support performance tuning across the fleet.
Platform Health and Deliverability Awareness
- Review internal code and configuration changes for their impact on sending infrastructure and reliability.
- Build and maintain tooling that tracks sending health and IP and domain reputation, and that supports the team's response to receiver and deliverability trends.
- Apply working knowledge of email authentication and receiver behavior (SPF, DKIM, DMARC, feedback loops, and bounce handling) to keep the platform configured correctly and delivering well.
Collaboration, AI, and Craft
- Partner with Senior and Staff Engineers and Product Managers to analyze infrastructure and reliability requirements, weighing scalability, security, cost, and performance as you design solutions.
- Help improve deployment methods, monitoring tools, and infrastructure automation practices across the team.
- Look for opportunities to automate operations and improve the platform with AI, and use AI and agent tooling to move faster while owning the quality of whatever those agents produce on your behalf.
- Document systems, runbooks, and operational workflows, act as a subject-matter expert, and mentor teammates to raise the whole team's capability.
- Communicate clearly and persuasively across global, cross-functional teams and time zones.