Director, Platform Software Engineering

Oracle
Nashville, TN, United States
6 days ago
Apply on eeho.fa.us2.oraclecloud.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Java (Programming Language) Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Microsoft Azure Cloud Computing Computer Clusters Software Debugging Linux Distributed Computing Environment Distributed Systems Fault Tolerance
+12 more
Network Control Operational Databases Oracle (Applications) Cloud Services Software Engineering System Software Google Cloud Autoscaling Kubernetes Azure AKS Machine Learning Operations Oracle Cloud Infrastructure

Job description

We are seeking a Director of Platform Software Engineering to lead teams responsible for core OKE platform capabilities. In this role, you will manage and develop engineering managers and senior technical leaders, set technical direction, and own delivery and operational outcomes for critical components of a highly available, globally distributed, 24x7 cloud service.

Your organization will advance Kubernetes cluster lifecycle management, orchestration, control plane reliability, scalability, performance, security, automation, and integration with OCI infrastructure. You will help evolve OKE to support larger clusters, more demanding enterprise workloads, and emerging AI and accelerated computing use cases.

This role requires deep Kubernetes knowledge, cloud infrastructure experience, strong distributed systems fundamentals, and a demonstrated ability to deliver through multiple engineering teams. You should be comfortable examining architecture and production behavior in detail, challenging technical assumptions, and guiding difficult decisions while empowering managers and engineers to own execution.

As a leader within OKE, you will partner with product management, senior architects, operations, and other OCI service teams to translate customer needs into a clear strategy and an achievable roadmap. Success requires sound judgment under ambiguity, disciplined execution, strong communication, and a commitment to developing people and improving the customer experience.

You will also guide the adoption of responsible AI-assisted and agentic engineering practices across design, implementation, testing, debugging, documentation, and operations. We expect you to help teams improve productivity while maintaining clear accountability for correctness, security, and production quality., * Lead, hire, coach, and develop engineering managers and software engineers. Build leadership capacity, establish clear expectations, and create a culture of ownership, collaboration, and technical excellence.

  • Define and execute the technical strategy and roadmap for core OKE platform capabilities, balancing customer needs, feature delivery, reliability, security, performance, and long-term maintainability.
  • Own delivery across multiple teams, including prioritization, staffing, dependencies, milestones, and risk management. Turn ambiguous requirements into clear plans and measurable outcomes.
  • Guide architecture and design for distributed systems that create, update, scale, repair, and operate Kubernetes clusters across OCI regions.
  • Provide technical leadership across Kubernetes control planes, controllers and operators, APIs, etcd, scheduling, autoscaling, container runtimes, and cluster and node lifecycle management.
  • Partner with OCI compute, networking, storage, identity, and security teams to deliver reliable integrations and resolve issues across service and organizational boundaries.
  • Own service health and operational outcomes for your organization’s components of OKE, including availability, capacity, performance, operational readiness, and customer escalations.
  • Lead effective incident response and ensure corrective actions address underlying causes and prevent recurring failures.
  • Establish rigorous standards for design reviews, testing, observability, production readiness, safe deployments, canary validation, upgrades, and rollback.
  • Drive automation that improves fleet health, detects failures earlier, accelerates diagnosis and recovery, and reduces manual operations.
  • Partner with product management and customers to understand workload requirements and use customer feedback and production data to guide investment decisions.
  • Prepare the platform for demanding AI/ML and GPU workloads, working across teams on scalability, orchestration, resource management, and infrastructure integration.
  • Introduce and scale responsible AI-assisted and agentic engineering workflows, measuring improvements in productivity and quality while maintaining security and human accountability.
  • Communicate strategy, delivery progress, service health, risks, and tradeoffs clearly to engineering teams, partners, and senior leadership.

Requirements

  • Extensive experience designing, building, and operating production software, including cloud infrastructure or distributed platform services.
  • Demonstrated success leading multiple engineering teams, managing engineering managers, developing technical leaders, and delivering complex initiatives across organizational boundaries.
  • Deep Kubernetes expertise and practical understanding of control plane architecture, cluster lifecycle, networking, storage, scalability, and production failure modes.
  • Strong distributed systems fundamentals, including availability, consistency, fault tolerance, performance, and operational tradeoffs.
  • Hands-on cloud infrastructure experience with OCI, AWS, Azure, GCP, or a comparable large-scale environment.
  • Strong software development background and the ability to guide design and implementation in Go and Java, supported by practical Linux, networking, and debugging knowledge.
  • Experience owning production service operations, incident response, safe change management, and sustained reliability improvements.
  • Strong judgment, communication, and execution skills, with the ability to lead through ambiguity and organizational change., * Experience building managed Kubernetes services or operating Kubernetes infrastructure at substantial scale.
  • Experience with Kubernetes networking and storage integrations, including CNI, CSI, Cilium, Calico, or equivalent technologies.
  • Experience with AI/ML infrastructure, GPU clusters, distributed training or inference, GPU scheduling, device plugins, or high-performance networking.
  • Experience improving engineering productivity through automation and responsible AI-assisted or agentic workflows.
  • Contributions to Kubernetes or related cloud native open-source projects.

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on eeho.fa.us2.oraclecloud.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

1:15 min

Deploying local container pods to Kubernetes clusters

Stevan Le Meur Stevan Le Meur · World Congress 2024

5:01 min

Container hosting options available on Microsoft Azure

Federico Fregosi · World Congress 2022

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

33 sec

Supporting NVIDIA GB200 GPUs on Kubernetes

Kevin Klues Kevin Klues · World Congress 2025

2:46 min

Evaluating managed container services and migration strategies

Adam Bien · World Congress 2021

Videos

See all

Related articles

See all