Principal Engineer, Core Infrastructure
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+21 more
Job description
- Lead the design and delivery of core distributed systems and cloud-platform capabilities.
- Drive complex technical work from design through implementation, rollout, and operational readiness.
- Work closely with architects and partner teams to improve the scale, resilience, performance, and simplicity of OCI services running on LWI.
- Solve difficult problems involving distributed state, failures, performance, security, lifecycle management, and automation.
- Raise the technical bar through design reviews, code reviews, mentoring, hiring, and hands-on technical leadership.
- Help define the engineering principles and long-term architecture of OCI’s next-generation service platform.
Where you can make an impact
Depending on your background and interests, you may work in one of the following LWI pillars:
- Compute: Build OCI’s next-generation multi-tenant compute platform, working across Rust, virtualization, QEMU, performance, and heterogeneous hardware.
- State: Build the storage layer for LWI-from fast temporary storage to durable volumes-and help shape the path toward serverless PostgreSQL on OCI.
- Identity & Security: Build hardware-rooted workload identity and security for multi-tenant cloud infrastructure, using PKI, confidential computing, enclaves, and attestation.
- Connectivity: Build the secure networking layer for LWI, including ingress, egress, eBPF-based firewalling, and high-performance internal load balancing.
- Stack Lifecycle and Infrastructure Management: Build the lifecycle and operations platform that deploys, upgrades, monitors, recovers, and securely images the entire LWI stack across OCI regions.
- Controllers: Build the Borg-style orchestration and management layer that schedules workloads, manages their lifecycle, and keeps OCI services running reliably at scale.
Requirements
- BS/MS in Computer Science, or equivalent practical experience.
- 6-10+ years of software engineering experience building highly available, scalable distributed systems or cloud infrastructure.
- Strong programming skills in Java, Go, Rust, C/C++, or a similar systems or backend language.
- Deep experience in one or more core infrastructure areas: distributed systems, cloud control planes, management planes, scheduling, virtualization, storage, networking, security, reliability, or platform automation.
- Ability to independently lead the design and delivery of complex services that must be correct, resilient, observable, and easy to operate.
- Strong understanding of failure handling, concurrency, asynchronous workflows, safe retries, and lifecycle management in distributed systems.
- Strong debugging and problem-solving skills, including the ability to drive resolution for complex production and performance issues.
- Experience collaborating across teams and turning broad platform goals into clear designs and executable plans.
- Ability to raise the technical bar through mentoring, design reviews, code reviews, and hands-on technical leadership.
Nice to have
- Experience building cloud platforms on OCI, AWS, Azure, or GCP.
- Experience with large-scale orchestration systems, schedulers, Kubernetes controllers, or Borg-style workload-management problems.
- Experience with Rust and systems programming, QEMU and virtualization, eBPF and Linux networking, distributed storage, PKI and confidential computing, image pipelines and software supply chain, fleet lifecycle management, or automated remediation.
- Experience working on multi-tenant workloads, high availability, zero-downtime changes, or global infrastructure.
Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.
About the company
Within OCI, the Technical Strategy & Oversight organization builds foundational systems for OCI’s most demanding services. One of its boldest initiatives is Autonomous OCI: a greenfield effort to build a cloud platform that can operate and scale with far less manual work.
Oracle Cloud Infrastructure (OCI) delivers mission-critical applications for enterprises around the world. OCI is expanding across public cloud, dedicated cloud, hybrid, multi-cloud, and edge deployments.
Within OCI, the Technical Strategy & Oversight organization builds foundational systems for OCI’s most demanding services. One of its boldest initiatives is Autonomous OCI: a greenfield effort to build a cloud platform that can operate and scale with far less manual work.
Today, every new region and service feature adds operational cost and complexity. Autonomous OCI is designed to change that by making services fast, resilient, and self-healing by default-helping OCI scale to thousands of regions without growing operations teams at the same rate.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
What Are The Top Skills Required For Azure Developers?
7 Cloud Computing Trends Coming in 2025 for Developers
Is Software Engineering Over-Saturated?