> Markdown version of [/jobs/ext/2170463-system-development-engineer-ii-ois-command-center](https://www.wearedevelopers.com/jobs/ext/2170463-system-development-engineer-ii-ois-command-center). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # System Development Engineer II, OIS Command Center - **Company:** Amazon.com, Inc. - **Location:** Arlington, VA, United States - **Salary:** $122,800.0 - $166,100.0 - **Contract:** Internship / Graduate position - **Skills:** Java (Programming Language), Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, C Sharp (Programming Language), C++ (Programming Language), Cyber Security, Computer Programming, Software Design Patterns, Linux, Distributed Systems, Amazon DynamoDB, Python (Programming Language), Machine Learning, Pattern Recognition, Windows PowerShell, Systems Development Life Cycle, Ruby, Software Engineering, Mobile Robots, Large Language Models, Information Technology, Serverless Computing, Amazon Redshift, Golang - **Published:** August 21, 2026 - **Apply:** https://dejobs.org/x/x/488D641FF6CD4EC9ABC972527A774CE4/job/ ## About the Role * 3+ years of non-internship professional software development experience * 1+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience * Experience in automating, deploying, and supporting large-scale infrastructure * Experience with Linux/Unix * Experience programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, Ruby * Bachelor's degree in computer science or equivalent, * 3+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience * 5+ years of non-internship professional software development experience * Experience building complex software systems that have been successfully delivered to customers, or experience with Machine Learning and Large Language Model fundamentals, including architecture, training/inference lifecycles, and optimization of model execution * Experience with AWS solutions such as EC2, DynamoDB, S3, and Redshift * Experience in security operations, risk management, and incident response * Experience with APIs and technical integrations * Experience with distributed systems at scale ## Description We're seeking an experienced Principal Technical Program Manager to lead the Join us in building Reflex, the agentic incident management platform for Amazon's fulfillment network. You'll design and deliver AI agents on Amazon Bedrock AgentCore that triage high-severity incidents, scribe live bridge calls in real time, draft stakeholder communications, and automate post-incident documentation and reporting, shifting incident management from a manual, pull-based model to an intelligent, push-based one. The OIS Command Center (OCC) is Amazon's 24/7 incident management function for high-severity incidents impacting fulfillment centers, delivery stations, and sortation centers worldwide, the infrastructure network that Amazon Robotics runs on. When this network degrades, robots stop and packages stop moving; OCC exists to make those minutes as short as possible. OCC manages roughly 1,500 high-severity incidents and triages some 14,000 alerts every year. Today, Incident Managers (IMs) continuously monitor signal feeds, engage resolver teams, and assemble a situational picture under time pressure before resolution work can even begin. Reflex changes that model fundamentally: agents watch the signals, assemble the context, and tell IMs when and how to engage, reserving human judgment for the decisions that actually need it. This is a builder role with an operational edge. Most of your time goes to designing, building, and operating Reflex agents and the platform beneath them: the agent runtime and tool orchestration on Amazon Bedrock AgentCore, the LLM evaluation framework that gates each agent's path from human-reviewed to autonomous, and the observability layer that keeps production agents accountable. You'll also periodically join live incident bridge calls in an Incident Manager capacity, staying close to the operational reality your software serves and turning what you learn on-call into what you build next. Your customers sit one Slack channel away, and you'll experience the impact of what you ship on the very next incident call., * Design, build, test, and operate AI agents and supporting services on AWS (Amazon Bedrock AgentCore, serverless compute, event-driven pipelines) that automate incident triage, call scribing, communications, post-incident documentation, and operational reporting * Own features end-to-end: from sitting with Incident Managers to understand the workflow, through design, implementation, evaluation, deployment, and production operation * Build the platform foundations that gate agent autonomy, including LLM output evaluation, monitoring and alerting for agents in production, and identity and access controls aligned with Amazon standards * Design the feedback loops through which agents learn from Incident Managers: capturing reviews, corrections, and approvals as evaluation signal, and turning resolved incidents into structured history that improves pattern matching, severity classification, and resolver routing over time * Integrate Reflex with the incident ecosystem: ticketing, chat, telemetry, detection feeds, and live call transcription. * Raise the bar on operational excellence, security, and quality for AI systems acting inside production incident workflows A day in the life You might start by reviewing overnight agent evaluation results and tuning a tool integration before shipping an improvement IMs see on the next incident. Later, you pair with an Incident Manager to observe how they used the scribing agent on a live call, turning their corrections into evaluation signal that moves the agent closer to autonomous posting. You also build the platform, design feedback loops, and integrate with ticketing, chat, and detection feeds. And periodically, you take a seat on a high-severity bridge call as an Incident Manager, because the best way to know what to automate next is to carry the workload firsthand. ## Related Videos - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Reliable scalability: How Amazon.com scales on AWS](https://www.wearedevelopers.com/videos/983-reliable-scalability-how-amazon-com-scales-on-aws) - [Fireside Chat with Werner Vogels, VP & CTO, Amazon.com & Daniel Gebler, CTO at Picnic](https://www.wearedevelopers.com/videos/1405-fireside-chat-with-werner-vogels-vp-cto-amazon-com-daniel-gebler-cto-at-picnic) - [Coroutine explained yet again 60 years later](https://www.wearedevelopers.com/videos/690-coroutine-explained-yet-again-60-years-later) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)