System DEvelopment Engineer, Automation, RTS-Solution Performance & Automation

Amazon.com, Inc.
North Reading, MA, United States
10 days ago
Apply on www.amazon.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Compensation
$129,200.0 - $174,800.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) Amazon Web Services C Sharp (Programming Language) C++ (Programming Language) Cloud Computing Cloud Engineering Code Review Continuous Integration Information Engineering Linux DevOps Distributed Systems
+13 more
Monitoring of Systems Python (Programming Language) Network Monitoring Systems Development Life Cycle Ruby Software Engineering Rust (Programming Language) Mobile Robots IT Architecture Software Troubleshooting Information Technology Software Version Control Golang

Job description

Amazon Robotics is seeking a System Development Engineer to join the Solution Performance & Automation (SPA) team in Robotics Technical Services (RTS) organization. SPA owns closed-loop network monitoring mechanisms: we define governed performance signals (what “good” looks like), standardize operational monitoring, and build scalable automation that converts signals into action (Andon, escalation, and operational workflows). You will build production monitoring and automation systems that make detection faster, response more consistent, and operations less manual. You will work within a well-understood strategy (Robotics Technical Services organization), and work in close collaboration with field teams to design and implement the monitoring systems, integrations, and mechanisms required to execute it at scale., Build and operate production monitoring services that evaluate robotics telemetry and performance signals for AR solutions, and trigger defined mechanisms (e.g., Andon, escalation, incident workflows) when conditions degrade. Implement signal evaluation and automation frameworks that are reusable across solutions; default to “build once, standardize, reuse everywhere” and document exceptions with rationale. Develop closed-loop operational automation including incident creation, routing, enrichment, and escalation integration with partner systems; reduce manual triage and repeated investigative effort. Improve signal quality and operator trust by reducing false positives/alert noise, tuning mechanisms, and instrumenting measurement for alert precision and operational outcomes. Build internal tooling that accelerates investigation and diagnosis for field/support stakeholders (explainability views, diagnostics tooling, guided triage workflows), grounded in “mechanisms over dashboards.” Own the end-to-end lifecycle of your services (design, implementation, testing, deployment, operations). When systems fail, ensure contributing causes are identified and eliminated with permanent fixes, not just mitigations. Be active in engineering and operational review mechanisms including code reviews, operational readiness reviews (ORRs), correction-of-errors (COEs), and post-incident analyses; use these mechanisms to drive resilience improvements and to coach peers. Balance constraints and explicitly manage short-term workarounds: avoid them where possible, replace them with long-term solutions, or escalate over-use when it creates systemic risk. Partner with Data Engineering on pipeline SLAs and reliability. Coordinate with solution/product engineering and field stakeholders through defined interfaces. Produce clear, accurate, inclusive documentation (technical runbooks, playbooks, operational procedures, and design notes) so systems and artifacts can be maintained and extended by engineers unfamiliar with them.

About the team Solution Performance & Automation (SPA) is RTS’s dedicated team for network monitoring mechanisms and automation across AR solutions. We provide a single, governed point of ownership for performance definitions, operational monitoring standards, and cross-team adoption. We build closed-loop monitoring mechanisms that define signals, automate Andon and escalation, and enable support scale through faster, informed decision-making.

Requirements

Bachelor’s degree in computer science or equivalent

  • Experience programming with at least one modern language such as Python, Ruby, Golang, Java, C++, C#, Rust
  • Experience in automating, deploying, and supporting large-scale infrastructure
  • Experience with Linux/Unix
  • Experience with version control systems and CI/CD pipeline implementation
  • Experience in automation or monitoring frameworks, deployment or development
  • 4+ years of systems design, software development, operations, automation, and process improvement experience
  • Experience troubleshooting and debugging technical systems, Experience with distributed systems at scale
  • Experience in any of the following: Cloud Architecture, Systems Design, Software Development, Infrastructure Architecture, Data Engineering or DevOps
  • Experience using data, reporting, or tools to measure performance and make adjustments accordingly
  • Experience in complex work environments, including (but not limited to robotics, automation, diagnostic and test equipment)
  • Experience working in a collaborative team environment to deliver high-quality design solutions
  • Experience building monitoring/observability systems, alerting mechanisms, and signal evaluation frameworks at scale
  • Experience with cloud infrastructure (AWS or equivalent), CI/CD, and operational readiness practices (ORRs/COEs/post-incident mechanisms)
  • Demonstrated ability to reuse/extend existing systems, make pragmatic tradeoffs, and reduce operational load through durable mechanisms

Benefits & conditions

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

USA, MA, North Reading - 129,200.00 - 174,800.00 USD annually

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.amazon.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

50 sec

Why developer happiness matters in web frameworks

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

1:06 min

Developer experience and project variety at scale

Alexandra Petri · World Congress 2023

3:30 min

Falling in love with Ruby and creating Basecamp

David Heinemeier Hansson David Heinemeier Hansson +1 · Coffee With Developers

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

Videos

See all

Related articles

See all