> Markdown version of [/jobs/ext/2689421-principal-software-development-engineer](https://www.wearedevelopers.com/jobs/ext/2689421-principal-software-development-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Software Development Engineer - **Company:** Expedia Inc. - **Location:** San Jose, CA, United States - **Salary:** $249,000.0 - $348,500.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Computing Platforms, Cloud Computing, Cloud Engineering, Computer Networks, Continuous Delivery, Continuous Integration, Programming Tools, Distributed Systems, Machine Learning, Performance Tuning, Prometheus, Service Discovery, Software Engineering, Systems Architecture, Rust (Programming Language), Pulumi, Google Cloud, Cloud Platform System, System Availability, Kubernetes, Api Design, Terraform, Dynatrace - **Published:** September 3, 2026 - **Apply:** https://www.careerbuilder.com/job-details/principal-software-development-engineer-cloud-platform-san-jose-ca--15d79ad3-bf46-4734-8975-05b45d39fc6c ## About the Role * Extensive professional software development experience designing, building, and operating large-scale, cloud-native distributed systems and platform services on Kubernetes. * Proven ownership of critical services or multi-service platforms, including responsibility for system design (LLD), API design, data modeling, deployment, and ongoing operational health. * Deep expertise with at least one major public cloud provider and core platform technologies (compute, networking, storage, service discovery, security, observability, and CI/CD). * Demonstrated ability to make high-impact architectural decisions, navigate complex trade-offs, and guide multiple teams toward coherent, long-term technical direction. * Familiarity with AI-driven systems, tools, or workflows and applying AI/ML concepts to real world products within cloud or platform environments. * Deep knowledge of observability patterns (OpenTelemetry, Prometheus, distributed tracing). * Expert-level understanding of Infrastructure as Code (Terraform, Pulumi) and CI/CD at scale. * Proficiency in Go, Rust, or similar languages used in modern platform engineering., * Track record of defining and evolving multi-year technical strategies for cloud and developer platform ecosystems, and successfully driving adoption of shared platforms across many teams. * Experience designing and operating highly available, globally distributed systems at internet scale, including capacity planning, performance optimization, and robust failure handling. * Safely integrates and operates AI/ML-enabled solutions that improve outcomes, such as intelligent routing, predictive scaling, or automated remediation embedded in platform services, with appropriate safeguards. * Advanced experience applying AI/ML techniques to cloud and platform problems (for example, cost optimization, anomaly detection, or performance tuning) and partnering with data/ML teams to productionize these capabilities. * A Systems Architect: You understand the deep plumbing of the cloud (AWS/GCP, K8s, networking). You think in terms of failure domains, latencies, and unit economics. * Reliability-First: You've carried a pager for global-scale systems. You have a healthy "paranoia" about state, consistency, and cascading failures. * Hands-on: You still love to build. You can prototype a complex infrastructure change in a weekend to prove it works, but you have the discipline to ensure it's production-grade before it ships., Amazon Web Services (AWS), Application Programming Interface (API), Architectural Services, Artificial Intelligence (AI), Budgeting, Capacity and Performance Management, Cloud Architecture, Cloud Computing, Computer Networks, Continuous Deployment/Delivery, Continuous Integration, Cost Control, Data Modeling, Distributed Computing, Economics, Ecosystems, Embedded Systems, Expense Tracking, GCP (Good Clinical Practices), Global Branding, High Availability, Incident Response, Leadership, Leading Edge Technology, Machine Tool, Pager, Performance Tuning/Optimization, Plumbing, Programming Tools, Public Cloud, Reporting Dashboards, Rust Programming Language, Software Design, Software Development, Software Engineering, System Architecture, Technical Strategy, Willing to Travel ## Description We are looking for a Principal Engineer to serve as the technical architect for our Cloud Platform organization which sits within our Technology division. As a Principal Engineer reporting to the VP of Cloud Platform, you will be the primary architect of our technical future. The Cloud Platform organization provides the secure, scalable cloud infrastructure, runtime platforms, and developer experience tooling that enable teams across Expedia Group to build, deploy, and operate high-quality, resilient software quickly and safely. We are seeing an explosion in code volume and service complexity. The goal for this role is to build a platform that can handle this growth without sacrificing reliability or skyrocketing our cloud bill. You'll be responsible for making sure our architecture is composable, our developer tools are agentic, our Kubernetes footprint is efficient, and our observability stack provides signals, not just noise. In this role, you will: * Lead Architectural Evolution: You'll own the move toward a Cell-Based Architecture. We need to move away from fragile, monolithic clusters and toward isolated, predictable failure domains that allow us to scale horizontally with confidence. * Modernize Kubernetes & Infrastructure: You'll define our K8s strategy, focusing on multi-cluster management, service mesh, and automated scaling. You need to ensure our "Golden Path" makes it easy for engineers to do the right thing by default. * Hardened Reliability & Observability: You will set the standards for SRE across the org. This means moving beyond basic dashboards to causal observability, automated incident response, and rigorous SLO/SLI management. You'll help us engineer out the root causes of systemic instability. * Optimize Cloud Economics: You'll lead our FinOps technical strategy. You need to build the tooling and visibility that allows us to understand cost-per-service and ensures our infrastructure spend is directly tied to business value. * Support the Developer Workflow: While we are embracing AI tools, your job is to build the underlying "agent-friendly" infrastructure. This includes standardized Dev Containers and ephemeral environments that allow for fast, isolated iteration without clobbering shared state. ## Related Videos - [Why segmenting your infrastructure into tiers makes your infrastructure design better](https://www.wearedevelopers.com/videos/1960-why-segmenting-your-infrastructure-into-tiers-makes-your-infrastructure-design-better) - [The Power of Purpose: Unlocking Potential and Innovation](https://www.wearedevelopers.com/videos/1110-the-power-of-purpose-unlocking-potential-and-innovation) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Terraform for Developers](https://www.wearedevelopers.com/videos/3-terraform-for-developers) - [Unleashing Potential Across Teams: The Power of Infrastructure as Code](https://www.wearedevelopers.com/videos/930-unleashing-potential-across-teams-the-power-of-infrastructure-as-code) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)