Sr Systems Dev Engineer, Amazon Leo OISL

Amazon.com, Inc.
Redmond, WA, United States
8 days ago
Apply on www.amazon.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Compensation
$151,200.0 - $204,600.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) Artificial Intelligence Amazon Web Services C++ (Programming Language) Computer Programming Extract Transform Load (ETL) Distributed Systems Embedded Software Apache Hive Python (Programming Language) Network Troubleshooting Network Monitoring
+16 more
Windows PowerShell Systems Development Life Cycle Cloud Services Prometheus Ruby Software Engineering SQL Databases Rust (Programming Language) Computer Network Operations Grafana Apache Spark Software Troubleshooting Performance Monitor Routing & Switching Data Pipelines Golang

Job description

We are seeking an exceptional Senior System Development Engineer to serve as the technical lead and senior individual contributor for our Network Operations team within the Optical Inter-Satellite Link (OISL) organization. This is a high-impact role where you will represent the entire Network Operations function of OISL - owning the strategy, tooling, automation, and operational excellence of a revolutionary laser communication network connecting thousands of satellites in space., This is a hybrid lead and senior individual contributor role. You will own the end-to-end Network Operations strategy for OISL, including defining what tools and services need to be built, driving automation of monitoring and incident response, and ensuring operational readiness as the constellation scales. You will be the senior technical voice for OISL Network Operations, partnering across the organization to drive improvements., Network Operations Strategy & Leadership

  • Own the OISL Network Operations strategy - tooling roadmap, automation priorities, and process maturity
  • Lead the NetOps on-call program: escalation design, automated ticketing, runbooks, and incident response
  • Drive operational cadence: metrics reviews, readiness assessments, and capacity planning

Automation & Tooling

  • Build scalable services that automate monitoring, fault detection, classification, and ticketing for the optical inter-satellite link network
  • Develop data pipelines for KPI tracking, anomaly detection, and proactive issue identification
  • Build observability platforms - dashboards, alerting, and real-time monitoring at scale
  • Apply ML/AI for fault correlation, root cause analysis, and predictive failure detection

Cross-Functional Collaboration

  • Partner with hardware, software, and ground network teams to drive system reliability improvements
  • Lead investigations into complex link failures and constellation-level fault patterns
  • Represent Network Operations in architecture reviews and program planning

Individual Contribution

  • Hands-on development of tools, services, and automation (Python, Spark SQL, cloud services)
  • Design testing strategies for distributed satellite network systems
  • Author technical specifications and operational documentation
  • Mentor team engineers

A day in the life You review the ticket queue and real-time satellite link health, spot degradation patterns, query telemetry, correlate with constellation changes, and engage and work with Hardware/Embedded Software Experts to root cause the issue. You lead cross-functional syncs on recurring failures, propose automated remediation workflows, build automated ticketing pipelines that eliminate manual triage, review teammates’ code, update on-call runbooks, and identify observability gaps to close.

About the team The OISL Network Operations team is responsible for the operational monitoring, troubleshooting, and reliability of 10,000+ optical laser links carrying customer traffic across a LEO satellite constellation. We build production-grade software services and tools to automate fault detection, classification, and resolution at scale. Our team operates at the intersection of network operations and software development - monitoring laser link health in real time, automating incident workflows, and driving cross-functional improvements with hardware and software engineering teams.

Requirements

6+ years of systems design, software development, operations, automation, and process improvement experience

  • Experience leading the design, automation, deployment, and support of large-scale infrastructure
  • Experience with PowerShell (preferred), Python, Ruby, or Java
  • Experience in Network Operations with proficiency in troubleshooting and debugging complex issues across Layer-1 (Optical Networking) or Layer-2/3 (Switching and Routing), 6+ years of experience in satellite network operations, telecommunications, or large-scale distributed network environments
  • Hands-on experience with network monitoring, availability/performance telemetry, and operational observability (Prometheus, Grafana, or similar)
  • Experience building data pipelines and analytics tooling (Spark, SQL, ETL workflows) for operational KPI tracking and anomaly detection
  • Experience troubleshooting Layer-1 and Layer-2/3 network issues in operational environments
  • Excellent communication skills with ability to influence cross-functional stakeholders and represent operations in engineering discussions
  • Programming proficiency in Python and at least one additional modern language (Go, Java, C++, Rust), with experience in cloud services (AWS) and distributed systems architecture

Benefits & conditions

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

USA, WA, REDMOND - 151,200.00 - 204,600.00 USD annually

About the company

Amazon LEO is an initiative to increase global broadband access through a constellation of 3,236 satellites in low Earth orbit (LEO). Its mission is to bring fast, affordable broadband to unserved and underserved communities around the world, helping close the digital divide for consumers, businesses, government agencies, and other organizations operating in places without reliable connectivity.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.amazon.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

50 sec

Why developer happiness matters in web frameworks

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

3:30 min

Falling in love with Ruby and creating Basecamp

David Heinemeier Hansson David Heinemeier Hansson +1 · Coffee With Developers

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all