Senior Software Engineer AI Infra

Oracle
Phoenix, AZ, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
4 years minimum
Working hours
Regular working hours
Job source

Tech stack

Board Bringup Java (Programming Language) Artificial Intelligence C++ (Programming Language) Cloud Computing Databases Data Centers Data Structures Software Debugging Linux Distributed Systems Network Interface Controllers
+23 more
Firmware Hardware Design Java Web Services Python (Programming Language) Network Architecture Network Protocols Object-Oriented Software Development Oracle (Applications) Performance Tuning Cloud Services Software Engineering TCP/IP Virtualization Technology AI Infrastructure Scripting Graphics Processing Unit (GPU) Linux Development Bare Metal Hardware Infrastructure Terraform Oracle Cloud Infrastructure Docker Microservices

Job description

Join Oracle Cloud Infrastructure’s Compute team to design, build, and scale the next generation of bare-metal provisioning systems powering millions of servers worldwide. As a senior engineer, you will develop highly reliable and secure infrastructure, tackle complex distributed systems challenges, and help deliver the foundation for OCI’s most performant compute services.

Oracle Cloud Infrastructure (OCI) is building the next generation of cloud services to support the world’s most demanding workloads. The Compute team is responsible for delivering bare-metal provisioning infrastructure that powers millions of servers and forms the foundation of OCI’s rapidly expanding AI infrastructure. The Compute Bare Metal Provisioning team owns the critical infrastructure responsible for automating the full server lifecycle from new platform shape (AMD/Intel/Arm/Nvidia) creation, hardware bring-up to customer-ready instance provisioning and firmware management.

The services operate at the intersection of bare metal hardware and full-stack orchestration frameworks, a unique combination where both distributed systems engineers and engineers with background in Linux and firmware are highly valued. The team interfaces directly with components like BMCs, NICs, SmartNICs, ILOMs, GPUs, and custom firmware stacks. The team builds high performance, scalable micro-services and tooling that provision, configure, secure, and validate server platforms across OCI’s massive fleet of Compute and GPU Infrastructure. You will partner closely across other teams in Compute, Networking, Security, Data center Engineering, and Hardware Development to ensure OCI can launch, scale, and maintain new server platforms with minimal operational overhead and high reliability. You will work directly with cutting edge GPU hardware and see the direct impact of your work on the business.

You are the builder here. You will be part of a team of really smart, motivated, and diverse people and given the autonomy and support to do your best work. It is a dynamic and flexible workplace where you’ll belong and be encouraged., As a Senior Member of Technical Staff, you will own the software design and development for major components of Oracle’s Cloud Infrastructure. You should be both a rock solid developer, driven problem solver and a distributed systems generalist and/or Linux developer with Systems experience able to dive deep, design, develop, operate, and debug any part of the stack and low level systems such as Linux, Docker, Java web services and Terraform, as well as design broad distributed system interactions. You should have a tenacious attitude to improve the status quo, independently seek out problems to solve and take action to deliver results wherever needed. You should value simplicity and scale, work comfortably in a collaborative, agile environment, and be excited to learn.

Requirements

  • 5-8 years’ experience delivering and operating large scale, highly available distributed systems, Linux development and Systems debugging.
  • Strong knowledge of Object Oriented programming such as C++ or Java, and experience with scripting languages such as Python.
  • Strong knowledge of data structures, algorithms, operating systems, and distributed systems fundamentals.
  • Experience with tools such as Terraform for Infrastructure as Code.
  • Working familiarity with networking protocols (TCP/IP, HTTP) and standard network architectures.
  • Strong understanding of databases, storage and distributed persistence technologies.
  • Strong troubleshooting and performance tuning skills.
  • Experience building multi-tenant, virtualized infrastructure a strong plus.

Duties and tasks are varied and complex needing independent judgment. Fully competent in own area of expertise. May have project lead role and or supervise lower level personnel. BS or MS degree or equivalent experience relevant to functional area. 4 years of software engineering or related experience.

About the company

If you are interested in building large-scale distributed infrastructure for the cloud, want to work on cutting edge GPU infrastructure and the latest Compute systems, have a knack for distributed systems and/or Linux development with Systems experience then this is your team! Oracle is aggressively investing in the Oracle Cloud to provide the broadest, most comprehensive cloud in the industry., Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

5:02 min

Mapping distributed compute paradigms to modern vehicles

Joachim Werner · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:50 min

Queues in TCP stacks and continuous network connections

Clemens Vasters Clemens Vasters · WWC 2022

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

Videos

See all

Related articles

See all