AKS / Platform Architect

Pioneer IT Systems LLC
United States
14 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Failover Role-Based Access Control Azure Active Directory Prometheus Data Logging Load Balancing Grafana Kubernetes Api Gateway Dynatrace

Job description

This team owns the container platform strategy for a large-scale enterprise environment. They’re building out multi-region Kubernetes infrastructure with a strong emphasis on automation, security, and developer self-service. This is a design-and-build role, not a maintenance role.

Requirements

  • 5+ years of Kubernetes and container platform architecture experience
  • Proven experience designing and deploying multi-region AKS clusters with failover and load balancing
  • Hands-on GitOps implementation using Flux, ArgoCD, or similar
  • Deep knowledge of Kubernetes ingress controllers and API gateway traffic management
  • Strong understanding of Kubernetes identity and access management, including Azure AD integration, RBAC, and pod identity
  • Experience implementing observability stacks: monitoring, logging, alerting, distributed tracing
  • Ability to travel for onsite work as needed (up to 25% annually) Nice-to-Have

  • CKA or CKAD certification
  • Experience with Prometheus, Grafana, ELK, or Jaeger specifically
  • Background mentoring platform engineering teams
  • Experience in a regulated or high-availability enterprise environment

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:01 min

Connecting frontend application performance to user retention and revenue

Dani Coll Dani Coll · World Congress 2025

2:59 min

Scaling clusters and handling automated replica failover

Jürgen Pilz · World Congress 2023

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:20 min

Architecting platform infrastructure for high availability and scale

Filippos Kyprianou +1 · World Congress 2023

12:08 min

Comparing Keptn orchestration capabilities against alternative software operators

Thomas Schütz · LIVE

Videos

See all

Related articles

See all