> Markdown version of [/jobs/ext/2419654-systems-development-engineer-gpu-ai-accelerator-servers-aws-hardware-engineering](https://www.wearedevelopers.com/jobs/ext/2419654-systems-development-engineer-gpu-ai-accelerator-servers-aws-hardware-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Systems Development Engineer, GPU & AI Accelerator Servers, AWS Hardware Engineering - **Company:** Amazon.com, Inc. - **Location:** Cupertino, CA, United States - **Salary:** $148,700.0 - $201,200.0 - **Contract:** Internship / Graduate position - **Skills:** Java (Programming Language), Agile Methodology, Artificial Intelligence, Amazon Web Services, Data Analysis, Systems Engineering, Build Automation, C Sharp (Programming Language), C++ (Programming Language), Cloud Computing, Computer Programming, Computer Engineering, Software Debugging, Software Design Patterns, Linux, Firmware, Python (Programming Language), PCI Express, Windows PowerShell, Scrum Methodology, Systems Development Life Cycle, Ruby, Software Engineering, Diagnostic Tools, Hardware Testing, Hardware Infrastructure, Network Server, Nvme, Golang - **Published:** August 31, 2026 - **Apply:** https://dejobs.org/x/x/6621DC5AB81747F5889E078CDE3B573F/job/ ## About the Role * 2+ years of non-internship professional software development experience * 1+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience * 3+ years of administrative experience in networking, storage systems, operating systems and hands-on systems engineering experience * Knowledge of systems engineering fundamentals (networking, storage, operating systems) * Experience programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, Ruby Preferred Qualifications * Experience with PowerShell (preferred), Python, Ruby, or Java * Experience working in an Agile environment using the Scrum methodology ## Description Fleet Health & Data Analysis * Analyze hardware failure patterns using fleet telemetry, system event logs, and datacenter tooling to identify root causes and quantify customer impact * Contribute to predictive failure detection using sensor data, error trending, and log correlation * Build and maintain operational dashboards and metrics for platform fleet health. * Build tooling to track component lifecycle (firmware versions, part revisions, supply chain status) across large-scale fleets Systems Development & Automation * Develop and maintain automation for hardware test, firmware qualification, and capacity recovery workflows * Develop diagnostic tools for Linux on ARM and x86 architectures * Debug and resolve Linux boot and runtime issues across processor architectures - PCIe, Power, NIC, NVMe, and GPU subsystems * Build automation solutions using Python, Java, or similar languages with focus on scalability and operational durability Cross-Team Collaboration * Collaborate with software, hardware, manufacturing, networking, and vendor teams to validate and qualify new compute solutions * Troubleshoot complex system-level issues in production environments, correlating across firmware, operating systems, drivers, and physical layers * Participate in sprint-based planning and oncall rotation for platform-level escalations A day in the life Some days you are deep in system event logs chasing a failure pattern across thousands of hosts; other days you are writing automation that eliminates a manual triage workflow entirely. You work with hardware engineers, firmware teams, datacenter operations, and vendor partners - driving quality and reliability from manufacturing through steady-state operations. Located in Cupertino, Seattle, or Denver, you work with global development teams on servers deployed in datacenters worldwide. About the team AWS Hardware Engineering designs and delivers next-generation cloud infrastructure - the servers, accelerators, and storage platforms that power AWS. Our team builds custom systems for AI training, inference, and compute workloads at global scale. We are directly responsible for launching and maintaining server hardware in the fleet, working across internal development teams, and design partners. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Fireside Chat with Werner Vogels, VP & CTO, Amazon.com & Daniel Gebler, CTO at Picnic](https://www.wearedevelopers.com/videos/1405-fireside-chat-with-werner-vogels-vp-cto-amazon-com-daniel-gebler-cto-at-picnic) - [Building Systems that Last](https://www.wearedevelopers.com/videos/1389-building-systems-that-last) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)