Principal Core Infrastructure Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+14 more
Job description
- Lead significant systems and initiatives from problem definition and design through implementation, rollout, adoption, and production validation.
- Translate scalability, security, reliability, and business requirements into clear technical designs and execution plans.
- Make sound tradeoffs involving availability, consistency, latency, throughput, durability, cost, and operational complexity.
- Design for partial failures, retries, duplicate requests, mixed-version deployments, dependency degradation, and regional disruption.
- Write and review secure, maintainable, well-tested Java code.
- Define service contracts, compatibility requirements, migration plans, validation strategies, and rollback criteria.
Scalability and Operational Excellence
- Establish capacity models, performance objectives, scaling strategies, and load-testing plans for high-throughput services.
- Design effective throttling, load shedding, backpressure, caching, concurrency, and failure-recovery mechanisms.
- Define useful service indicators, objectives, metrics, alarms, dashboards, runbooks, and deployment safeguards.
- Lead complex incident investigations and convert recurring failures or manual procedures into automation and preventive engineering improvements.
- Serve as a technical escalation point for problems that cross application, infrastructure, network, or organizational boundaries.
Technical Leadership
- Provide architectural direction in one or more critical areas such as routing, authentication, private connectivity, runtime performance, observability, or deployment infrastructure.
- Decompose broad initiatives so multiple engineers can own meaningful work while maintaining architectural consistency.
- Mentor engineers through design, code review, delivery, and incident response.
- Raise engineering quality through reusable systems, tools, standards, and operational practices.
- Contribute to hiring and help identify architectural investments, platform gaps, and reliability risks for the team roadmap.
Cross-Team Execution
- Align SPLAT and partner teams on technical decisions, responsibilities, dependencies, and rollout plans.
- Communicate complex designs, tradeoffs, risks, and progress clearly to engineers and leaders.
- Make progress under ambiguity by separating facts, assumptions, reversible decisions, and external dependencies.
- Adjust direction when production evidence or new technical information invalidates earlier assumptions.
- Use modern development and AI-assisted tools responsibly to improve engineering quality and productivity.
Requirements
- Bachelor’s degree in Computer Science, Computer Engineering, or a related field, or equivalent practical experience.
- 8+ years of experience designing, building, and operating production backend or platform services.
- Strong development experience in Java or another modern object-oriented language.
- Strong understanding of distributed systems, concurrency, fault tolerance, and production operations.
- Experience leading substantial technical initiatives across multiple engineers or teams.
- Demonstrated ability to diagnose complex production issues and improve service reliability.
- Strong written and verbal communication skills., * Experience with high-throughput, low-latency services, HTTP, networking, proxies, or API gateways.
- Experience with authentication, authorization, TLS, certificates, private connectivity, or multi-tenant security.
- Experience with throttling, load shedding, caching, and capacity planning.
- Experience with cloud infrastructure, Kubernetes, infrastructure as code, CI/CD, and deployment automation.
- Experience improving a broader engineering organization through mentoring, shared tooling, or technical standards., * Software Design and Development
- Backend Programming Languages
- Distributed Systems
- System Design
Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.
About the company
Lead the architecture, development, and operation of highly available Java platform services that securely route hundreds of billions of API requests each month for more than 300 OCI control planes.
Oracle Cloud Infrastructure builds and operates large-scale cloud services in a distributed, multi-tenant environment. The SPLAT team owns critical platform services that provide secure API routing, service registration, traffic management, private connectivity, authentication, and operational controls for OCI services.
SPLAT sits in the request path for more than 300 OCI control planes and processes hundreds of billions of API requests each month. Our engineering challenges span high-throughput Java services, distributed systems, networking, security, observability, capacity management, deployment automation, and production reliability.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
93 Java Interview Questions You Should Prepare For
What Are The Top Skills Required For Azure Developers?
What is Software Engineering in the Age of AI?