> Markdown version of [/jobs/ext/2079095-senior-support-engineer](https://www.wearedevelopers.com/jobs/ext/2079095-senior-support-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Support Engineer - **Company:** Oracle - **Location:** Harwood, MD, United States - **Experience:** Expert - **Salary:** $51,900.0 - $132,400.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, Oracle (Applications), AI Infrastructure, Reliability of Systems, AI Platforms, Cts+, Oracle Cloud Infrastructure - **Published:** August 16, 2026 - **Apply:** https://www.careerjet.com/job/us4100eb9740a7e64742c783e892dc7c65/eaa ## About the Role Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements. Range and benefit information provided in this posting are specific to the stated locations only ## Description The Hardware Systems Engineer is responsible for diagnosing, testing, repairing, and validating next-generation AI compute hardware supporting Oracle Cloud Infrastructure (OCI). As a core member of the Oracle Repair Center, this role ensures Compute hardware is rapidly returned to service while helping improve hardware reliability, diagnosability, and operational efficiency across the repair ecosystem. The position plays an important role in sustaining repair velocity and supporting the broader scale-up of Oracle's AI infrastructure. Working closely with Development, Manufacturing, Supply Chain, and Global Product Engineering, the Hardware Systems Engineer performs advanced hardware troubleshooting, root cause analysis, and physical repair of AI infrastructure, with a particular focus on NVIDIA and AMD compute hardware and future AI platforms. The role includes functional validation of Compute Trays (CTs) using standardized repair and test processes, execution of repair validation through dedicated CDU test infrastructure, and completion of engineering observation reports and repair documentation for manufacturing, quality, and engineering teams. In addition, the engineer serves as a technical resource during coolant leak events, supports continuous improvement efforts, and may participate in weekend coverage and on-call rotations for critical events., * Diagnose, troubleshoot, repair, and validate NVIDIA, AMD, and future-generation AI compute hardware, including Compute Trays (CTs), utilizing standardized repair and validation processes. * Execute functional validation using dedicated CDU test infrastructure to ensure repaired hardware meets manufacturing and operational quality standards before returning to production. * Perform advanced hardware failure analysis, root cause investigation, and engineering observation reporting to improve product reliability, diagnosability, and repair effectiveness. * Partner with Development, Manufacturing, Supply Chain, Quality, and Global Product Engineering to investigate systemic hardware issues and implement corrective actions. * Serve as the primary technical resource during coolant leak events, coordinating with field technicians to assess hardware condition, support containment activities, and document engineering observations. * Participate in new product introductions (NPI) and future platform readiness activities, helping develop repair procedures, validation processes, and operational standards for next-generation AI platforms. * Provide technical support through the Global Operations Technical Support ticket queue, assisting with complex hardware issues and remote troubleshooting across Go Big repair centers. * Support launch activities for new Go Big repair centers, including knowledge transfer, technician training, operational readiness, and travel to domestic or international sites as business needs require. * Participate in weekend coverage and on-call rotations to support critical business operations and high-priority customer events. Core Responsibilities * Deliver safe, high-quality repairs while consistently meeting established turnaround time, throughput, and quality objectives. * Maintain detailed repair documentation, engineering observations, service records, and knowledge articles to support continuous improvement and engineering feedback. * Identify recurring hardware failure trends and recommend improvements to repair processes, tooling, diagnostics, automation, and operational workflows. * Collaborate across Repair Operations, Development, Manufacturing, Supply Chain, Quality, and Engineering to resolve technical issues and improve product serviceability. * Build and maintain technical expertise on Oracle AI infrastructure, including current and future GPU platforms, repair methodologies, diagnostics, and test systems. * Contribute to a culture of operational excellence through knowledge sharing, mentoring, cross-functional collaboration, and continuous learning. * Effectively prioritize work across multiple repair priorities while maintaining flexibility to respond to critical customer issues, surge events, and changing business demands. * Support regional and global repair-center scalability by helping standardize repair processes, operational best practices, and training materials across the Oracle Repair Centers. * Demonstrate sound technical judgment, ownership, and accountability while working independently and collaboratively in a fast-paced, mission-critical repair environment. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) - [This App Reached 10,000 Users in One Week. Here's How.](https://www.wearedevelopers.com/videos/100329-this-app-reached-10-000-users-in-one-week-here-s-how) - [AI-Augmented DevOps with Platform Engineering](https://www.wearedevelopers.com/videos/1614-ai-augmented-devops-with-platform-engineering) - [The 2026 Talent Pivot: Essential Trends for an Evolving Workforce](https://www.wearedevelopers.com/videos/100098-the-2026-talent-pivot-essential-trends-for-an-evolving-workforce) - [How to build a sovereign European AI compute infrastructure](https://www.wearedevelopers.com/videos/1102-how-to-build-a-sovereign-european-ai-compute-infrastructure) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)