Director, Cloud Engineering
About the role
The Director of Cloud Engineering is a senior engineering leader accountable for the strategy, architecture, and operational excellence of the company's multi-cloud platform. Reporting to the Sr. Director of IT Operations, this role owns the end-to-end cloud estate across Google Cloud Platform (GCP) and Microsoft Azure, setting the technical direction for cloud-native infrastructure, platform engineering, and developer experience. You will build and lead a high-performing distributed team of cloud engineers, SREs, and platform engineers while partnering closely with product, security, data, and application engineering to deliver a reliable, secure, and cost-efficient cloud foundation. This is a hands-on leadership role that balances deep technical credibility with the strategic thinking and organizational influence expected at the Director level.
Responsibilities
- Define and own the multi-cloud architecture strategy across GCP and Azure, aligning platform capabilities with business objectives, scalability requirements, and total cost of ownership targets.
- Lead architectural decisions for cloud-native workloads, including microservices, containerization (Kubernetes/GKE/AKS), serverless functions, and event-driven patterns.
- Drive the evolution of the Internal Developer Platform (IDP), enabling self-service infrastructure provisioning, golden-path templates, and standardized deployment pipelines.
- Establish and enforce cloud governance frameworks, including resource hierarchy, landing zones, policy-as-code, and guardrails across both cloud providers.
- Evaluate emerging technologies—AI/ML infrastructure, LLMOps tooling, edge compute—and make build-vs-buy recommendations grounded in cost, risk, and strategic fit.
- Own Infrastructure as Code (IaC) standards and practices; ensure all infrastructure changes flow through version-controlled, peer-reviewed pipelines.
- Oversee GitOps workflows and CI/CD pipeline infrastructure, partnering with application teams to reduce deployment lead time and increase release frequency.
- Direct SRE practices across the cloud estate: define and track SLIs/SLOs, drive blameless post-mortems, and lead reliability improvements through systematic error-budget management.
- Own observability strategy—logging, metrics, tracing, and alerting—across cloud environments, ensuring engineers have actionable insight into system health and performance.
- Lead disaster recovery design, runbook development, and regular DR/BCP testing exercises; drive remediation of identified gaps to closure.
- Administer and maintain all Microsoft licensing through NCE and MPSA agreements.
- Maintain network architecture and connectivity across cloud environments, data centers, and supported locations, including VPN, private connectivity (Interconnect/ExpressRoute), and DNS.
- Partner with the Security team to implement Zero Trust network architecture, enforce least-privilege IAM, and operationalize Cloud Security Posture Management (CSPM) tooling.
- Ensure cloud environments meet regulatory and compliance requirements; support internal and external audits by providing architecture documentation, access logs, and control evidence.
- Define and maintain cloud security baselines, secrets management practices, and data encryption standards across the shared responsibility model.
- Drive incident response processes for cloud infrastructure events, ensuring timely escalation, containment, and root-cause resolution.
- Own the cloud infrastructure budget; implement FinOps practices including cost allocation tagging, chargeback/showback reporting, reserved capacity planning, and rightsizing recommendations.
- Establish cloud cost visibility and governance mechanisms so engineering teams can make cost-aware architecture decisions in real time.
- Negotiate vendor agreements and manage strategic relationships with GCP, Azure, and key third-party tooling providers.
- Evaluate, integrate, and champion the responsible use of AI tools across cloud engineering and platform teams, setting standards for AI-driven productivity and innovation.
- Treat AI as a productivity multiplier, not a replacement — ensuring engineers maintain sound engineering judgment and that AI-generated code and configuration meet the same quality, security, and reliability standards as hand-written work.
- Lead, mentor, and grow a distributed team of cloud engineers, platform engineers, and SREs (onshore and offshore); establish clear career ladders and development paths.
- Build a team culture grounded in psychological safety, continuous learning, and engineering excellence; champion internal tech talks, documentation habits, and knowledge sharing.
- Manage headcount planning, recruiting, and onboarding; partner with HR and technical leads to define roles and evaluate candidates.
- Deliver timely, constructive performance reviews; proactively manage performance issues and recognize high-impact contributions.
- Serve as a technical escalation path and executive sponsor for major infrastructure initiatives; represent the cloud engineering team in IT governance and steering forums.
Qualifications
- 15+ years in IT/infrastructure engineering, with at least 8 years focused on public cloud platforms (GCP and/or Azure).
- 5+ years in a senior engineering leadership role (Director, Principal, or Staff Engineer equivalent) managing teams of 8 or more engineers across onshore and offshore locations.
- Demonstrated hands-on proficiency with both GCP and Azure services, including but not limited to: GKE, Cloud Run, Cloud Armor, BigQuery, VPC Service Controls (GCP) and AKS, Azure Policy, Defender for Cloud, ExpressRoute, and Entra ID (Azure).
- Proven track record delivering large-scale cloud migrations, re-architecture projects, or greenfield platform builds in a regulated or enterprise environment.
- Deep expertise with Infrastructure as Code (Terraform required; Pulumi or CDK a plus) and GitOps tooling (ArgoCD, Flux, or equivalent).
- Experience designing and operating CI/CD pipelines at enterprise scale (GitHub Actions, Cloud Build, Azure DevOps, or equivalent).
- Direct experience with FinOps practices: cloud cost allocation, tagging governance, reserved instance/committed-use optimization, and budget reporting.
- Familiarity with SRE principles: SLI/SLO frameworks, error budgets, chaos engineering, and reliability-focused incident management.
- Bachelor’s degree in Computer Science, Information Systems, or a closely related engineering discipline, or equivalent professional experience.
- At least one of: Google Cloud Professional Cloud Architect, Google Cloud Professional DevOps Engineer, Microsoft Azure Solutions Architect Expert (AZ-305), or Microsoft Azure DevOps Engineer Expert (AZ-400).
Preferred qualifications
- Experience building or maturing an Internal Developer Platform (IDP) using tools such as Backstage, Crossplane, or similar platform-engineering frameworks.
- Exposure to AI/ML infrastructure on GCP (Vertex AI) or Azure (Azure AI Studio, Azure Machine Learning) and an understanding of LLMOps patterns such as model serving, inference optimization, and vector database infrastructure.
- Working knowledge of service mesh technologies (Istio, Linkerd) and zero-trust network segmentation at the workload level.
- Experience with cloud-native observability stacks (Prometheus/Grafana, Datadog, Dynatrace, or Google Cloud Operations Suite).
- Familiarity with enterprise IT governance frameworks such as ITIL v4, COBIT 2019, or TOGAF; direct experience supporting SOC 2, ISO 27001, or HIPAA/PCI-DSS compliance programs.
- Experience in financial services, healthcare, or another regulated industry where cloud security posture and audit readiness are operationally critical.
- Additional certifications: Certified Kubernetes Administrator (CKA), HashiCorp Terraform Associate, Google Cloud Professional Security Engineer, or AWS Solutions Architect (for multi-cloud context).
Pay
- $161,000-$200,000 per year