Principal Software Development Engineer - Cloud Platform
About the Role
Our Technology Team partners with teams across Expedia Group to create innovative products, services, and tools to deliver high-quality experiences for travelers, partners, and our employees. We are looking for a Principal Engineer to serve as the technical architect for our Cloud Platform organization within our Technology division. As a Principal Engineer reporting to the VP of Cloud Platform, you will be the primary architect of our technical future.
Responsibilities
- Lead Architectural Evolution: Own the move toward a Cell-Based Architecture, transitioning from monolithic clusters to isolated, predictable failure domains for horizontal scaling.
- Modernize Kubernetes & Infrastructure: Define our Kubernetes strategy, focusing on multi-cluster management, service mesh, and automated scaling, ensuring a "Golden Path" for engineers.
- Hardened Reliability & Observability: Set standards for SRE, moving beyond basic dashboards to causal observability, automated incident response, and rigorous SLO/SLI management.
- Optimize Cloud Economics: Lead FinOps technical strategy, building tooling for cost-per-service visibility and tying infrastructure spend to business value.
- Support the Developer Workflow: Build "agent-friendly" infrastructure, including standardized Dev Containers and ephemeral environments for fast, isolated iteration.
Qualifications
- Extensive experience designing, building, and operating large-scale, cloud-native distributed systems and platform services on Kubernetes.
- Proven ownership of critical services or multi-service platforms, including system design, API design, data modeling, deployment, and operational health.
- Deep expertise with at least one major public cloud provider and core platform technologies (compute, networking, storage, service discovery, security, observability, and CI/CD).
- Demonstrated ability to make high-impact architectural decisions and guide multiple teams toward coherent, long-term technical direction.
- Familiarity with AI-driven systems and applying AI/ML concepts to cloud or platform environments.
- Deep knowledge of observability patterns (OpenTelemetry, Prometheus, distributed tracing).
- Expert-level understanding of Infrastructure as Code (Terraform, Pulumi) and CI/CD at scale.
- Proficiency in Go, Rust, or similar languages used in modern platform engineering.
Preferred Qualifications
- Track record of defining and evolving multi-year technical strategies for cloud and developer platform ecosystems.
- Experience designing and operating highly available, globally distributed systems at internet scale.
- Advanced experience applying AI/ML techniques to cloud and platform problems (e.g., cost optimization, anomaly detection).
- Systems architecture expertise in cloud plumbing (AWS/GCP, K8s, networking) with a focus on failure domains, latencies, and unit economics.
- Reliability-first mindset with experience carrying a pager for global-scale systems.
- Hands-on prototyping and production-grade implementation skills.
Pay
The total cash range for this position in San Jose is $249,000.00 to $348,500.00. Employees in this role have the potential to increase their pay up to $398,500.00, based on ongoing, demonstrated, and sustained performance. Starting pay will vary based on location, budget, and individual experience.
Benefits
Expedia Group offers benefits and perks including medical, dental, and vision coverage, paid time off, wellness and travel reimbursement, travel discounts, and International Airlines Travel Agent Network (IATAN) membership.