Systems Support Engineer, Reverse Recycle...

Amazon.com, Inc.
Greencastle, PA, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
1 year minimum
Compensation
$82,400.0 - $144,100.0
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) PHP (Programming Language) Computing Platforms Systems Engineering Bash Shell Configuration Management Data Centers DevOps Perl (Programming Language) Firmware Python (Programming Language) Linux System Administration
+14 more
Networking Basics Network Administration Performance Tuning Queue Management Systems Reliability Engineering Ansible Ruby Software Engineering AI Infrastructure Hardware Testing System-level Testing Information Technology Hardware Infrastructure Puppet

Job description

Shape the future of AI infrastructure! Join RRL and develop system-level test solutions for our innovative ML acceleration hardware, deployed across our global server fleet.

We’re seeking highly skilled and motivated Product Test Engineers to join our RRL test & repair operation. In this role, you’ll be at the forefront of validating and ensuring that test infrastructure is deployed and operating at required scale. You’ll be responsible for designing and implementing comprehensive system-level test strategies that cover the full spectrum of our ML acceleration products, from individual components to fully integrated systems. This position requires a unique blend of hardware knowledge, software expertise, and systems thinking, as you’ll be working at the intersection of custom silicon, complex firmware, and high-performance ML workloads.

You’ll collaborate closely with cross-functional teams including hardware designers, software engineers, and operations specialists to develop robust test solutions that can scale to meet the demands of RRL’s global infrastructure. Your work will be crucial in identifying and resolving integration issues, optimizing system performance, and ultimately ensuring that our ML acceleration products meet the highest standards of reliability and efficiency in real-world data center environments. If you’re passionate about pushing the boundaries of ML hardware testing and have a knack for solving complex system-level challenges, we want you on our team.

Key job responsibilities

Provide frontline technical support to manufacturing and operations teams by troubleshooting complex tester, hardware, and system-level failures impacting production throughput

Own daily ticket queue management, prioritizing operationally critical issues to ensure timely resolution and minimal business disruption

Perform deep technical troubleshooting of production test systems, identifying root causes across hardware, software, networking, and infrastructure dependencies

Diagnose, repair, and restore test equipment and validation environments to maintain operational readiness and maximize uptime

Partner closely with Operations, Engineering, and cross-functional stakeholders to drive issue resolution and improve overall test process stability

Execute structured root cause analysis and implement corrective actions to prevent recurring equipment or process failures

Support system bring-up, tester validation, and configuration activities for new and existing production environments

Develop and maintain troubleshooting documentation, standard work, and knowledge-sharing mechanisms to improve team efficiency and issue resolution consistency

Identify opportunities to simplify support workflows, improve response times, and enhance operational effectiveness through process improvements and automation where applicable

Monitor tester health, system performance, and operational trends to proactively identify risks and escalate issues before business impact occurs

Requirements

2+ years of Linux systems administration and/or development experience

  • 1+ years of work in at least two of these languages: Python, Java, Perl, PHP, Ruby or Bash/Shell experience

  • Bachelor’s degree in Computer Science or other technical degree or related experience

  • Knowledge of networking fundamentals

  • Experience working in a 24/7 production environment

  • Experience in Linux systems administration and/or development

  • Experience working in at least two of these languages: Python, Java, Perl, PHP, Ruby or Bash/Shell

Preferred Qualifications

  • 2+ years of site reliability engineering (SRE), systems engineering, systems administration, DevOps, security administration, or network administration experience

  • 1+ years of building scripts, tooling, and automation for large-scale computing environments experience

  • Knowledge of configuration management systems, such as Puppet, Chef, Ansible, or related systems

  • Experience in site reliability engineering (SRE), systems engineering, systems administration, DevOps, security administration, or network administration

  • Experience in network capture and systems troubleshooting

  • Experience building scripts, tooling, and automation for large-scale computing environments

Benefits & conditions

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits .

USA, PA, Greencastle - 82,400.00 - 144,100.00 USD annually

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on juju.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

50 sec

Why developer happiness matters in web frameworks

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

2:26 min

Understanding Puppeteer and its underlying architectural design

Miki Lombardi · JS Congress

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:30 min

Falling in love with Ruby and creating Basecamp

David Heinemeier Hansson David Heinemeier Hansson +1 · Coffee With Developers

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

Videos

See all

Related articles

See all