Senior GCP Operations Architect
540 is seeking a Senior GCP Operations Architect to support a mission-critical aviation platform that enables the safe and efficient operation of the nation's airspace. The platform modernizes how time-sensitive safety information is managed and distributed through secure, near-real-time data exchange.
Location: Remote within the continental United States
Citizenship & Clearance Requirement: Candidates must be U.S. Citizens and currently hold or be eligible to obtain and maintain an FAA Public Trust
Education Requirement: Bachelor's Degree in Computer Science or related field (preferred)
About the role
Working alongside Google technical advisors and customer operations teams, you will provide architectural guidance and day-two operational expertise across GCP infrastructure, security, resiliency, automation, and DevSecOps. This highly visible role requires a proactive technical leader who can resolve complex issues, strengthen operational readiness, and help ensure the reliability and availability of a platform that directly supports aviation safety.
Responsibilities
- Provide expert guidance on the architecture and operation of GCP networking, compute, security, storage, and supporting platform services
- Partner with customer operations teams to identify and introduce new tools, operational patterns, and cloud best practices that improve team productivity, platform reliability, and operational efficiency
- Serve as a primary point of contact for infrastructure and operations questions and technical escalations, providing Tier 3 troubleshooting across cloud infrastructure, Terraform deployments, and DevSecOps pipelines
- Quickly develop expertise and take ownership of unfamiliar technical or operational areas, providing guidance to customer and delivery teams as new challenges emerge
- Participate in Agile delivery ceremonies and technical forums, including daily stand-ups, sprint planning and reviews, technical syncs, scrum-of-scrums, operational reviews, postmortems, and technical retrospectives
- Review system metrics and logs to identify performance bottlenecks, stability concerns, and emerging risks; support root cause analysis and define sustainable, long-term remediation actions
- Assess and strengthen day-two operational readiness, including troubleshooting procedures, escalation paths, support-team preparedness, backup and recovery strategies, disaster recovery protocols, and high-availability configurations
- Provide guidance on Terraform module development, state management, infrastructure automation, and CI/CD integration, and review GitOps and DevSecOps workflows to troubleshoot critical pipeline and deployment issues
- Support release-management activities, including the application and rollback of database releases
- Review and guide the development of runbooks, SOPs, troubleshooting materials, and other operational artifacts to ensure they are accurate, complete, consistent, and aligned with the current architecture and recommended practices
- Review the Technical Design Document every two weeks, advise customer stakeholders on required revisions and document governance, and develop supporting materials for technical and architecture reviews
- Conduct weekly office hours for Infrastructure, DevSecOps, and Release Management processes
- Engage with the customer's Change Control Board and Architecture Review Board to evaluate proposed configuration and architectural changes
- Research relevant GCP services and architectural components to inform technical recommendations and customer decision-making
- Maintain accurate and detailed activity, status, and task information in Jira
- Reduce technical and operational ambiguity by clarifying requirements, establishing a path forward, and maintaining momentum when requirements or direction are not fully defined
- Demonstrate proactive ownership and confident technical leadership across the broader project, helping maintain architectural cohesion, operational standards, and delivery quality while representing Google and 540 in customer-facing engagements
Requirements
- 8+ years of relevant experience
- Extensive experience designing, operating, or supporting production environments in Google Cloud
- Strong knowledge of GCP networking, compute, storage, identity, access management, and core security services
- Experience translating cloud architecture into practical day-two operations and support processes
- Hands-on experience with Terraform, including reusable modules, remote state management, deployment troubleshooting, and CI/CD integration
- Experience supporting GitOps, CI/CD, or DevSecOps workflows in production environments
- Experience with infrastructure monitoring, logging, system-health analysis, and operational risk identification
- Strong understanding of high availability, disaster recovery, backup and recovery, and platform resiliency
- Experience conducting root cause analysis and developing long-term corrective and preventive actions
- Ability to troubleshoot complex cloud infrastructure and deployment issues at a Tier 3 level
- Experience creating and reviewing technical design documents, runbooks, SOPs, troubleshooting guides, and operational procedures
- Experience supporting release-management and rollback processes
- Working knowledge of Agile delivery practices and task-management tools such as Jira
- Strong customer-facing communication, facilitation, and technical advisory skills
- Ability to communicate architectural concepts and operational recommendations to both technical teams and customer leadership
- Proactive ownership mindset with the confidence to identify risks, recommend improvements, and drive issues toward resolution
Nice to have
- Google Cloud professional-level certification
- Experience supporting mission-critical government or regulated cloud environments
- Experience with Apigee architecture, deployment, or operations
- Experience participating in formal Change Control Board or Architecture Review Board processes
- Experience supporting database release and rollback activities
- Familiarity with site reliability engineering and production operations practices
- Experience supporting large-scale modernization or cloud-transformation programs
- Familiarity with aviation, transportation, or other high-availability operational systems
Benefits
- Flexible PTO + all Federal holidays off
- Health, dental and vision insurance plans
- Flexible Spending Account (FSA)
- 401k with employer match
- Company-sponsored life insurance, short- and long-term disability
- Professional development (training, certifications, conferences)
- Paid cloud developer accounts
- Referral bonuses
- HQ office perks (parking / metro reimbursement, nitro coffee & lunches)
- Annual social events (540 Week, hackathon, charity golf tournament, etc.)
- Access to 540's Washington Capitals & Nationals tickets