Principal Engineer, OCI Object Storage

Oracle
Nashville, TN, United States
6 days ago
Apply on eeho.fa.us2.oraclecloud.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Java (Programming Language) Application Programming Interfaces (APIs) Artificial Intelligence Big Data Cloud Computing Computer Programming Data Security Software Debugging Software Design Patterns Disaster Recovery File Systems Distributed Data Store
+25 more
Distributed Systems Memory Management Fault Tolerance Java Web Services Meta-Data Management Network Control Object-Oriented Software Development Oracle (Applications) Queueing Systems E2e Testing Cloud Services Software Engineering Software Systems Encapsulation (Networking) Data Logging Multithreading System Availability Grafana Event Driven Architecture Containerization Low Latency Deployment Automation Restful APIs Oracle Cloud Infrastructure Programming Languages

Job description

We are seeking a Principal Core Infrastructure Engineer to help design, build, and operate the foundational systems behind a highly available, durable, and globally distributed object storage service. In this senior individual-contributor role, you will solve complex distributed-systems problems at massive scale and set technical direction across critical storage-platform capabilities.

You will work with engineers across storage, networking, security, control plane, and observability to deliver resilient systems that customers can trust with their most important data.

What You’ll Do

  • Architect and deliver core infrastructure for object storage, including metadata, data placement, replication, lifecycle management, and durability workflows.
  • Lead the design of distributed systems that provide strong availability, consistency, performance, and operational simplicity at scale.
  • Improve service reliability through fault isolation, automated remediation, disaster recovery design, capacity planning, and rigorous operational practices.
  • Drive technical strategy for complex, cross-team initiatives; influence architecture and execution beyond your immediate team.
  • Investigate and resolve challenging production issues, using deep systems expertise to prevent recurrence.
  • Build tooling, automation, and observability that make the service easier to operate, diagnose, and evolve.
  • Partner with security and compliance teams to embed secure-by-design practices into storage infrastructure.
  • Establish engineering standards through design reviews, technical mentorship, and clear written communication.

Requirements

  • Strong software engineering experience in Java or a similar programming language, including designing, developing, testing, debugging, and maintaining production-quality services.
  • Deep understanding of object-oriented programming principles, including abstraction, encapsulation, inheritance, polymorphism, composition, and appropriate application of common design patterns.
  • Experience designing clean, extensible APIs and service interfaces, with an emphasis on maintainability, testability, backward compatibility, and operational safety.
  • Demonstrated experience designing and building distributed systems, including services that operate across multiple hosts, availability domains, or regions.
  • Strong knowledge of distributed-systems concepts such as replication, consistency, consensus, partition tolerance, leader election, idempotency, retries, failure handling, and eventual consistency.
  • Experience designing systems for high availability, scalability, fault tolerance, and data durability.
  • Experience with concurrent and multithreaded programming, performance analysis, memory management, and diagnosing latency or throughput bottlenecks in Java services.
  • Proficiency with unit, integration, and end-to-end testing practices, including designing tests for failure scenarios and distributed-system edge cases.
  • Experience using observability tools and practices, including metrics, logging, tracing, alerting, and production debugging.
  • Ability to write clear technical design documents, evaluate architectural tradeoffs, and lead design reviews for complex services.
  • Strong collaboration skills and experience partnering with engineering, security, networking, and operations teams.
  • Proven ability to mentor engineers, raise engineering standards, and provide technical leadership without direct people-management responsibility.

Preferred Qualifications

  • Experience with object storage systems, distributed databases, file systems, or other large-scale data platforms.
  • Knowledge of storage-system concepts such as metadata management, object lifecycle operations, data placement, replication or erasure coding, and recovery from hardware or infrastructure failures.
  • Experience designing systems that span multiple regions or availability domains.
  • Experience operating services with demanding availability, durability, latency, or throughput requirements.
  • Familiarity with RESTful APIs, asynchronous processing, message queues, and event-driven architectures.
  • Experience with containerized environments, cloud infrastructure, and automated deployment pipelines.
  • Knowledge of security practices for cloud services, including authentication, authorization, encryption, and secure data handling.

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on eeho.fa.us2.oraclecloud.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:07 min

Architecting the availability stack with Prometheus and Grafana

Gabriel Labachelerie · World Congress 2023

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

1:46 min

Understanding how Pathway ensures low latency data processing

Bobur Umurzokov · LIVE

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all