> Markdown version of [/jobs/ext/2652001-principal-systems-software-engineer-gpu-platform-systems](https://www.wearedevelopers.com/jobs/ext/2652001-principal-systems-software-engineer-gpu-platform-systems). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Systems Software Engineer - GPU Platform Systems - **Company:** Oracle - **Location:** Sacramento, CA, United States - **Experience:** Expert - **Salary:** $114,600.0 - $234,600.0 - **Contract:** Permanent contract - **Skills:** Board Bringup, C (Programming Language), Adobe InDesign, Systems Engineering, Build Automation, Bash Shell, Booting (BIOS), C++ (Programming Language), Cloud Computing, Code Review, Communications Protocols, Computer Programming, Data Centers, Software Debugging, Linux, Embedded Software, Emulators, Fault Tolerance, Firmware, Field-Programmable Gate Array (FPGA), Hardware Interface Design, Joint Test Action (IEEE Standards), Python (Programming Language), Software Maintenance, Remote Access Service, Software Deployment, Software Engineering, Systems Integration, Software Technical Review, Data Logging, Diagnostic Tools, Scripting, Application Specific Integrated Circuits, Computer Equipment, Information Technology, Hardware Infrastructure, Oracle Cloud Infrastructure, Server Operating Systems & Platforms, Hardware Debugging - **Published:** August 16, 2026 - **Apply:** https://www.techcareers.com/job.asp?id=3355969949&tx=JK1509FFK&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role This role is well suited for an engineer with deep experience in computer systems, embedded or platform firmware, and hardware/software integration who enjoys solving difficult problems that cross traditional engineering boundaries., * 6+ years of programming and/or scripting experience, with relevant languages such as C, C++, Python, or Bash. * Strong systems-software, embedded-software, or firmware development experience. * Experience integrating software or firmware with complex hardware systems. * Strong understanding of computer hardware fundamentals and the interaction between hardware, firmware, operating systems, and software. * Demonstrated ability to diagnose complex problems that cross hardware and software boundaries. * Experience with software development practices including design, implementation, debugging, code review, testing, automation, and quality assurance. * Experience working in Linux/Unix-based development or systems environments. * Ability to independently own technically complex projects and collaborate across multiple engineering organizations. Preferred Qualifications: Experience in one or more of the following areas is highly desirable * BMC, OpenBMC, service processors, or server platform-management technologies. * Server, GPU, accelerator, or other complex compute-platform firmware. * GPU or server RAS, telemetry, fault management, observability, or serviceability. * Board, device, or system bring-up. * Hardware debugging using JTAG, logic analyzers, emulators, schematics, or related diagnostic tools. * Firmware lifecycle management and secure firmware-update mechanisms. * Power management, power control/capping, thermal management, or platform telemetry. * Hardware interfaces and low-level communication protocols. * CPU, GPU, SoC, ASIC, or FPGA-based systems. * ARM, AMD, Intel, NVIDIA, or similarly complex compute platforms. * Automation and diagnostics for server provisioning, validation, or fleet operations. * Large-scale cloud or data-center infrastructure. * Technical leadership, mentoring, architecture/design ownership, and cross-functional engineering coordination. ## Description Oracle Cloud Infrastructure (OCI) is seeking a Principal Systems Software Engineer to help build and evolve the low-level systems software and platform-management capabilities that power next-generation GPU infrastructure. This is a hands-on systems engineering role at the intersection of software, firmware, hardware, and large-scale cloud infrastructure. You will work on complex GPU and server platforms, owning critical capabilities spanning BMC/service-processor software, platform management, firmware lifecycle, reliability and serviceability, telemetry, power management, hardware bring-up, and fleet operations. You will work closely with silicon, firmware, hardware, compute, and fleet engineering teams, as well as technology and manufacturing partners, to bring new platforms from initial hardware enablement through production deployment and ongoing operation at cloud scale., BMC & Platform-Management Systems * Design, develop, and maintain complex BMC, service-processor, and platform-management software for GPU and server infrastructure. * Develop capabilities for host management, baseboard management, telemetry, platform monitoring, power control, and system diagnostics. * Design and implement secure, reliable firmware-management and update workflows across complex, multi-vendor platforms. * Define robust interfaces and integration contracts between platform firmware, hardware components, operating systems, drivers, and higher-level infrastructure services. Systems & Firmware Development * Design and implement low-level systems software and firmware using technologies such as C, C++, Python, and Bash. * Develop maintainable software for managing, monitoring, diagnosing, and provisioning server and GPU systems. * Build automation and tooling that improves platform provisioning, onboarding, validation, diagnostics, and fleet operations. * Develop and debug advanced platform capabilities involving areas such as RAS, telemetry, power management, high-speed I/O, and chipset/SoC services. * Conduct design and code reviews and help establish strong engineering practices for maintainability, testing, observability, security, and reliability. Hardware Bring-Up & Cross-Layer Debugging * Play a leading role in initial board, device, and platform bring-up for new GPU and server systems. * Diagnose difficult failures spanning hardware, firmware, bootloaders, operating systems, drivers, and platform services. * Use hardware and software diagnostic techniques-including logs, schematics, JTAG, logic analyzers, emulators, and platform instrumentation-to isolate root causes. * Partner with silicon, board, firmware, and manufacturing teams to validate end-to-end platform behavior and resolve integration issues. * Turn complex or recurring failures into durable engineering fixes, improved diagnostics, automation, and preventive controls. GPU Reliability, Serviceability & Operations * Develop reliability, availability, and serviceability (RAS) capabilities for large-scale GPU and server environments. * Improve telemetry, fault detection, logging, observability, and diagnostics used to identify and resolve platform issues. * Develop and improve power-control and power-capping capabilities and related platform instrumentation. * Design systems with fleet-scale reliability, fault tolerance, secure firmware lifecycle, and operational serviceability in mind. * Support difficult platform incidents and escalations and help translate field findings into long-term product and engineering improvements. Technical Leadership * Own technically complex and sometimes ambiguous areas from architecture and design through implementation, validation, and deployment. * Drive technical decisions and establish clear interfaces across teams responsible for different layers of the platform. * Lead deep technical investigations and help teams reach evidence-based root causes for difficult system failures. * Raise engineering standards through architecture and design reviews, code reviews, testing practices, automation, and diagnostic tooling. * Mentor and provide technical guidance to engineers while remaining actively involved in design, coding, bring-up, and debugging. * Collaborate effectively across software, firmware, hardware, silicon, compute, fleet, support, and external partner organizations. ## Related Videos - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Playing Pong on a shoulder press machine](https://www.wearedevelopers.com/videos/100140-playing-pong-on-a-shoulder-press-machine) - [Thinking Differently - How to Make Money from Cyber Attacks & Cheats](https://www.wearedevelopers.com/videos/745-thinking-differently-how-to-make-money-from-cyber-attacks-cheats) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Agent Smith Gets Hardware: Autonomous IoT Hacking From Debug Port to Cloud API](https://www.wearedevelopers.com/videos/100258-agent-smith-gets-hardware-autonomous-iot-hacking-from-debug-port-to-cloud-api) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Hard?](https://www.wearedevelopers.com/magazine/448-is-software-engineering-hard) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs)