Lead Principal Core Infrastructure Engineer

Oracle
Santa Clara, CA, United States
5 days ago

Role details

Contract type
Permanent contract
Employment type
Part-time (≤ 32 hours)
Experience level
Expert
Experience required
10 years minimum
Compensation
$146,300.0 - $306,400.0
Working hours
Regular working hours

Tech stack

C (Programming Language) HTML Java (Programming Language) Agile Methodology Artificial Intelligence Big Data C Sharp (Programming Language) C++ (Programming Language) Cloud Computing Program Optimization Cyber Security Concurrent Computing
+37 more
Data Integrity Data Security Distributed Systems Failover Fault Tolerance Formal Verification Home Automation Python (Programming Language) Network Security Transport Layer Machine Learning Network Architecture Network Protocols Oracle (Applications) Performance Tuning Scrum Methodology Systems Development Life Cycle Cloud Services Requirements Management Service Pack Software Engineering Systems Architecture System Testing TCP/IP Transmission Control Protocol (TCP) Data Processing Scripting System Availability Concurrency Software Security Software Troubleshooting Reliability of Systems Infrastructure as Code (IaC) Information Technology Database Replication Golang Programming Languages

Job description

Mentors teams and leads the architecture of highly scalable, interdependent distributed systems. Identifies and removes performance/scalability bottlenecks for hyper-scale workloads; defines scalability requirements with stakeholders; and designs elastic, high-impact systems while advancing innovation in data plane platforms. Engineers and oversees fault-tolerant, in-service-upgradable designs; optimizes resilience mechanisms (load-shedding, throttling, rate-limiting); and sets SLO-aligned durability and availability standards across dependent services. Establishes KPIs and advanced telemetry; applies formal verification for complex features; and develops robust replication/synchronization strategies. Advises and leads resolution of complex production issues, sets operational readiness and SOP standards, and directs incident response and RCAs. Architects advanced security controls, drives remediation and compliance, and delivers enterprise-level automation (IaC) and change strategies enabling safe, automated patching, updates, and rollbacks.

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs., System Design & Architecture - System Scalability:

-Mentor the team in the architecture and design of highly scalable, interdependent distributed systems, ensuring horizontal and vertical scalability and overall performance, including leveraging distributed state management tools.

-Lead the identification of performance and scalability bottlenecks and recommend solutions to optimize code and/or systems for large-scale data processing and high-throughput requirements to improve performance for hyper-scale systems.

-Lead collaboration with stakeholders to define system scalability requirements, ensuring the defined requirements meet customer expectations.

-Leverage deep expertise to design high-impact, interdependent systems to scale with elasticity (e.g., effectively scaling both up and down).

-Drive innovation in the use of data plane platforms.

-Evaluate whether systems are meeting nonfunctional scalability requirements, and proactively anticipate growing business needs within the business unit.

System Design & Architecture - System Reliability Design:

-Design and oversee the implementation of fault-tolerant, interdependent systems capable of withstanding in-service updates by implementing sophisticated redundancy, replication, and automatic failover capabilities.

-Lead the design and implementation of systems that effectively handle service disruptions (e.g., network partitions) by prioritizing consistency, availability, or partition tolerance.

-Guide the optimization of advanced mechanisms to handle network unreliability, including load-shedding, throttling, and rate-limiting.

-Design interdependent systems that are durable and adhere to service level objectives (SLOs), driving standards for availability and durability of other computing services within the organization

System Design & Architecture - System Reliability Performance:

-Define key performance indicators (KPIs) and telemetry to identify risks, gaps, or cyclical dependencies in running, interdependent systems.

-Drive the creation and customization of highly complex dashboards, telemetry systems, and alerting mechanisms, proactively ensuring system health and reliability.

System Design & Architecture - Correctness / Availability:

-Maintain expertise in industry standards for verifying correctness and apply existing techniques to interdependent systems.

-Formally verify complex features (e.g., via TLA+) to ensure system design correctness for various interdependent systems.

-Develop advanced strategies for data replication and synchronization, ensuring robust data integrity and availability

Compliance & Security:

-Architect advanced security measures to protect data and applications in multi-tenant environments, and lead initiatives to enhance data and application protection.

-Guide the execution of comprehensive remediation plans to address identified security vulnerabilities.

-Ensure cloud infrastructure is in compliance with industry standards and regulations, and guide documentation efforts across projects.

Automation & Change Management:

-Develop enterprise-level automation tools and strategies (e.g., Infrastructure as Code (IaC)) and oversee their implementation.

-Drive alignment of change management plans and organizational initiatives for patching, updating, and rolling back applications, and design interdependent systems to allow for automation of these processes.

Requirements

  • Bachelor’s degree in Computer Science or equivalent proven experience
  • 10+ years of experience building and operating large scale, highly available, cloud based distributed systems
  • Specialist skill in a modern programming language such as Java, C, C++, C#, Go, or Python, with proficiency in additional languages preferred
  • Validated understanding of operating system fundamentals
  • Strong understanding of data models and distributed persistence technologies
  • Thorough understanding of the latest security principles, techniques, and protocols
  • Strong troubleshooting and performance tuning skills
  • Proficiency in network, distributed, asynchronous, and concurrent programming
  • Knowledge of professional software engineering standard methodologies for the full software development process
  • Experience building and operating scalable infrastructure software or distributed systems
  • Proven track record to achieve stretch goals in a highly innovative and fast-paced environment
  • Passion for technical leadership and mentoring
  • Strong verbal and written communication skills
  • Strong analytical skills, with excellent problem-solving abilities

Preferred Qualifications

  • Experience in Agile/SCRUM enterprise-scale software development
  • Experience building automated network and security solutions
  • Knowledge of Machine Learning fundamentals
  • Working familiarity with networking protocols (TCP/IP, HTTP) and standard network architectures
  • Working familiarity with storage principles, protocols and practices
  • Working familiarity with building secure software using modern security principles, Accidental Death and Dismemberment (AD&D), Agile Programming Methodologies, Analysis Skills, Applications Security, Artificial Intelligence (AI), Automation, Business Growth, C Programming Language, C++ Programming Language, Change Management, Cloud Computing, Communication Skills, Computer Science, Computer Security, Computer Services, Concurrency, Concurrent Programming Language Family, Customer Relations, Data Modeling, Data Processing, Data Quality, Database Replication, Dental Insurance, Distributed Computing, Documentation, Embedded Systems, Establish Priorities, Failover, Financial Planning, Flexible Spending Accounts, Go Programming Language (Golang), HTTP (HyperText Transport Protocol), Healthcare, High Availability, High Throughput, Home Automation, Identify Issues, Incident Response, Industry Standards, Information/Data Security (InfoSec), Infrastructure Software, Java, Large-Scale Systems, Legal, Life Insurance, Machine Learning, Maintain Compliance, Mentoring, Microsoft C# (C Sharp), Network Architecture/Engineering, Network Protocols, Network Security, Occupational Health, Operating Systems, Oracle, Organizational Development/Management, Performance Management, Performance Metrics, Performance Tuning/Optimization, Presentation/Verbal Skills, Problem Solving Skills, Process Improvement, Programming Languages, Property Insurance, Python Programming/Scripting Language, Regulations, Regulatory Compliance, Replication and Remote Mirroring, Reporting Dashboards, Requirements Management, Risk Analysis, Scalable System Development, Scrum Project Management and Software Development, Security Architecture, Software Design, Software Development, Software Engineering, Software Patches, Standard Operating Procedures (SOP), Stock Purchase Plans, Strategic Planning, System Architecture, System Validation, Systems Reliability, Systems Scalability, TCP/IP (Transmission Control Protocol/Internet Protocol), Technical Leadership, Telemetry, Test Requirements, Traffic Shaping, Vision Plan, Writing Skills

Benefits & conditions

Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Disclaimer:

Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.

Range and benefit information provided in this posting are specific to the stated locations only

US: Hiring Range in USD from: $146,300 to $306,400 per annum. May be eligible for bonus, equity, and compensation deferral.

Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle’’s differing products, industries and lines of business.

Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

Oracle US offers a comprehensive benefits package which includes the following:

  1. Medical, dental, and vision insurance, including expert medical opinion
  2. Short term disability and long term disability
  3. Life insurance and AD&D
  4. Supplemental life insurance (Employee/Spouse/Child)
  5. Health care and dependent care Flexible Spending Accounts
  6. Pre-tax commuter and parking benefits
  7. 401(k) Savings and Investment Plan with company match
  8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
  9. 11 paid holidays
  10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
  11. Paid parental leave
  12. Adoption assistance
  13. Employee Stock Purchase Plan
  14. Financial planning and group legal
  15. Voluntary benefits including auto, homeowner and pet insurance

About the company

For over three decades, Oracle has been the center of innovation for business software birthplace of the first commercially available relational database, the first suite of internet-based applications, and the next-generation enterprise-computing platform, Oracle Fusion. Today, Oracle provides the world’s most complete, open, and integrated business software and hardware systems, with more than 370,000 customers including - 100 of the Fortune 100 - representing a variety of sizes and industries in more than 145 countries around the globe. And Oracle’s 110,000 global employees - including 30,000 developers working full-time on Oracle products -are critical to that success. Oracle Supports Workforce Diversity

Company Size: 10,000 employees or more

Industry: Computer/IT Services

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

5:02 min

Mapping distributed compute paradigms to modern vehicles

Joachim Werner · LIVE

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:21 min

Projecting external HTML content using default and named slots

Rowdy Rabouw Rowdy Rabouw · WWC 2022

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:50 min

Queues in TCP stacks and continuous network connections

Clemens Vasters Clemens Vasters · WWC 2022

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · WWC Europe 2026

Videos

See all

Related articles

See all