engineer lead, Cloud Foundation Services Platform (Kubernetes)

Starbucks
Seattle, WA, United States
1 day ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Working hours
Regular working hours
Job source

Tech stack

Agile Methodology Amazon Web Services Computing Platforms Systems Engineering Microsoft Azure Cloud Computing Cloud Engineering Cyber Security System Configuration Continuous Delivery Continuous Integration Software Debugging
+26 more
DevOps Distributed Systems Python (Programming Language) Key Management Octopus Deploy Open Source Technology Platform as a Service (PAAS) Reliability Engineering Zero Trust Network Access Service Discovery Software Engineering Usage Analysis Data Logging Google Cloud Istio System Availability Git Containerization Kubernetes Infrastructure Automation Frameworks Information Technology Bicep Azure AKS Azure Service Fabric Terraform Oracle Cloud Infrastructure

Job description

We are seeking a Cloud Engineer - Lead with deep expertise in Azure Kubernetes Service (AKS) and cloud-native platform engineering to design, build, and operate a secure, scalable, and cost-efficient Kubernetes-based platform on Microsoft Azure. This role will serve as a hands-on technical leader responsible for enabling development teams through reliable platforms, automation, and modern DevOps and GitOps practices., Technical Leadership & Collaboration

  • Communicate complex platform and Kubernetes architecture decisions clearly to both technical and non-technical stakeholders
  • Establish strong cross-functional partnerships with application, security, networking, and business teams
  • Act as a technical leader and mentor for Kubernetes, AKS, and Azure platform engineering best practices
  • Partner with technology vendors and open-source communities to deliver against business and platform objectives

AKS & Platform Engineering

  • Design, implement, and operate Azure Kubernetes Service (AKS) platforms at scale
  • Build and maintain a secure, multi-tenant Kubernetes platform with a focus on:
  • Reliability, performance, and scalability
  • Environment consistency across dev, test, and production Leverage Azure-native services and PaaS offerings to support application and platform needs

  • Own Kubernetes cluster lifecycle management, upgrades, capacity planning, and operational health

Kubernetes Ecosystem & Service Mesh

  • Design and operate Kubernetes-native solutions using:
  • Argo CD / Argo Workflows for GitOps and continuous delivery
  • Istio (or similar service mesh) for traffic management, security, and observability
  • Implement best practices around ingress, egress, service discovery, and policy enforcement
  • Enable secure communication, workload identity, and secrets management across the platform

Automation, Infrastructure as Code & GitOps

  • Drive an automation-first approach using:
  • Terraform, Azure Bicep, and Crossplane for infrastructure and platform provisioning
  • Python for scripting, tooling, and automation
  • Enable CI/CD and GitOps workflows using Azure DevOps, Git-based pipelines, and modern delivery patterns
  • Engineer standardized, repeatable build and release processes for platform and application teams

Security, Compliance & Cost Optimization

  • Design platforms that are secure by default, following Zero Trust and least-privilege principles
  • Ensure all implementations align with Information Security policies and compliance requirements (e.g., PCI)
  • Continuously evaluate and optimize platform cost using cloud-native tooling and usage analysis
  • Implement system configurations and baselines that support secure software development and operational best practices

Observability & Operations

  • Implement deep telemetry, logging, and monitoring across Kubernetes and Azure PaaS platforms
  • Deploy and maintain observability solutions to support performance, reliability, and troubleshooting
  • Ensure high availability and operational continuity for mission-critical services
  • Support and operate 24x7 production environments, enabling automated recovery wherever possible

Requirements

The ideal candidate is an AKS and Kubernetes expert with strong working knowledge of Azure PaaS services, service mesh technologies, and continuous delivery systems, and who brings a strong focus on security, reliability, and cost optimization., * Bachelor’s degree in Computer Science, Engineering, or a related field (or equivalent professional experience)

  • 8-10 years of professional industry experience in cloud, platform, or software engineering
  • 2+ years of experience leading teams or serving as a technical lead for groups of four or more engineers, Cloud & Platform Expertise
  • Strong hands-on experience with Microsoft Azure (mandatory)
  • Strong working experience with Amazon Web Services (AWS) (Good to have)
  • Exposure to Google Cloud Platform (Google Cloud Platform) and/or Oracle Cloud Infrastructure (OCI) is a plus
  • Deep hands-on experience with Azure Kubernetes Service (AKS) and Kubernetes platform operations

Kubernetes & Cloud-Native Technologies

  • Expert-level Kubernetes knowledge, including:
  • Cluster architecture, scheduling, scaling, and upgrades
  • Ingress/egress, service discovery, and workload isolation
  • Experience with service mesh technologies such as Istio
  • Strong experience with Argo CD / GitOps-based delivery models
  • Familiarity with Azure PaaS services (e.g., storage, messaging, identity, secrets)

Automation & DevOps

Strong hands-on experience with:

  • Terraform, Azure Bicep, and Crossplane
  • Python for automation and tooling
  • CI/CD experience using Azure DevOps, GitOps workflows, and modern DevOps tools
  • Strong understanding of DevOps and Agile principles
  • Observability & Reliability
  • Experience implementing application and infrastructure logging and monitoring solutions
  • Proven experience operating and supporting high-scale, highly available platforms
  • Ability to troubleshoot complex distributed systems and performance issues, * 8+ years of experience in systems engineering, platform engineering, or site reliability engineering
  • Experience with large-scale distributed systems and cloud-native architectures
  • Proven ability to debug, optimize, and automate complex systems
  • Strong understanding of security best practices in Kubernetes and cloud platforms
  • Knowledge of regulatory requirements such as SOX, PCI, HIPAA, and data protection standards
  • Ability to translate technical insights into actionable business and platform recommendations

Benefits & conditions

As a Starbucks partner, you (and your family) will have access to medical, dental, vision, basic and supplemental life insurance, and other voluntary insurance benefits. Partners have access to short-term and long-term disability, paid parental leave, family expansion reimbursement, paid vacation from date of hire*, sick time (accrued at 1 hour for every 25 hours worked), eight paid holidays, and two personal days per year. Starbucks also offers eligible partners participation in a 401(k) retirement plan with employer match, a discounted company stock program (S.I.P.), Starbucks equity program (Bean Stock), incentivized emergency savings, and financial well-being tools. Additionally, Starbucks offers 100% upfront tuition coverage for a first-time bachelor’s degree through Arizona State University’s online program via the Starbucks College Achievement Plan, student loan management resources, and access to other educational opportunities. You will also have access to backup care and DACA reimbursement. Starbucks will comply with any applicable state and local laws regarding employee leave benefits, including, but not limited to providing time off pursuant to the Colorado Healthy Families and Workplaces Act, and in accordance with its plans and policies. This list is subject to change depending on collective bargaining in locations where partners have a certified bargaining representative. For additional information regarding partner perks and more detailed information about benefits, go to starbucksbenefits.com.

*If you are working in CA, CO, IL, LA, ME, MA, NE, ND or RI, you will accrue vacation up to a maximum of 120 hours (190 in CA) for roles below director and 200 hours (316 in CA) for roles at director or above. For roles in other states, you will be granted vacation time starting at 120 hours annually for roles below director and 200 hours annually for roles director and above.

The actual base pay offered to the successful candidate will be based on multiple factors, including but not limited to job-related knowledge/skills, experience, geographical location, and internal equity. At Starbucks, it is not typical for an individual to be hired at the high end of the range for their role, and compensation decisions are dependent upon the facts and circumstances of each position and candidate.

If you live in the greater Seattle area, we offer a flexible workplace that allows for hybrid work. Partners can work remotely up to one day per week.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:56 min

Provisioning a secure container infrastructure with Bicep

Matthias Falkenberg +1 · WWC 2022

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · WWC Europe 2026

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

4:09 min

Selecting infrastructure tools and determining proper abstraction layers

Alayshia Knighten Alayshia Knighten · WWC 2024

7:15 min

Installing Istio programmatically with bash scripts

Thomas Südbröcker · LIVE

Videos

See all

Related articles

See all