Senior/Lead GCP Platform Engineer
Capgemini · Atlanta, GA · 4 days ago
Engineering$105k–$115k/yrFull-time
Key Responsibilities
- Design, build, and maintain enterprise-scale cloud platform services on Google Cloud Platform
- Develop reusable platform capabilities, standards, and automation frameworks
- Implement Infrastructure as Code (IaC) using Terraform
- Support cloud governance, IAM, networking, security, and compliance requirements
- Drive platform reliability, scalability, and operational excellence
- Design and implement centralized observability architecture across multiple GCP projects and environments
- Build enterprise monitoring, logging, tracing, and alerting solutions using GCP native and third-party observability tools
- Centralize logs, metrics, traces, dashboards, and alerts across development, test, and production environments
- Define observability standards, SLOs, alerting strategies, and operational runbooks
- Create executive, operational, and application health dashboards
- Establish monitoring coverage for cloud infrastructure
- Develop solutions for automated alert management, reporting, and operational analytics
- Partner with development and operations teams to improve system reliability
- Conduct root cause analysis and incident investigations
- Optimize monitoring coverage while reducing alert fatigue
- Support production operations and platform upgrades
- Drive continuous improvement of operational processes and platform observability
Required Technical Skills
- Google Cloud Platform: Strong hands-on experience with Google Cloud Monitoring and Google Cloud Logging
- Managed Service for Prometheus, Cloud Trace, Cloud Profiler, and Cloud Audit Logs
- Experience with Pub/Sub, Cloud Storage, BigQuery, Cloud Functions, Cloud Run, GKE (Google Kubernetes Engine)
- Experience with IAM, Shared VPC, Organization level monitoring configurations, Observability Platforms
- DevOps Automation, Terraform, Git/GitHub, CI/CD pipelines, Kubernetes
- Experience with one or more: Datadog, Splunk, Dynatrace, Grafana, Prometheus, OpenTelemetry
Required Experience
- 5 years of cloud platform engineering experience
- 3 years of hands-on Google Cloud Platform experience
- 3 years of Python development and automation experience
- Experience implementing centralized observability solutions across multiple cloud accounts/projects
- Experience with enterprise monitoring, logging, alerting, and operational analytics
- Experience supporting Kubernetes-based platforms
Preferred Qualifications
- Experience designing organization-wide observability frameworks
- Experience implementing OpenTelemetry standards
- Knowledge of Site Reliability Engineering (SRE) practices
- Experience with FinOps and cloud cost optimization
- Experience supporting data platforms and AI/ML workloads
- Google Cloud Professional Cloud Architect certification
- Google Cloud Professional DevOps Engineer certification
Key Deliverables
- Enterprise-wide observability architecture across GCP projects
- Centralized monitoring and logging platform
- Automated project onboarding to observability services
- Standardized dashboards, alerts, and operational metrics
- Improved platform reliability and reduced incident response times
- Python-based automation framework for cloud operations
Pay
$105,000 - $115,000
Benefits
- Paid time off based on employee grade (A-F): Vacation 12-25 days depending on grade, Company paid holidays, Personal Days, Sick Leave
- Medical, dental, and vision coverage
- Retirement savings plans (401(k) in the U.S., RRSP in Canada)
- Life and disability insurance
- Employee assistance programs
- Other benefits as provided by local policy and eligibility