> Markdown version of [/jobs/ext/2873789-staff-or-principal-engineer-distributed-systems](https://www.wearedevelopers.com/jobs/ext/2873789-staff-or-principal-engineer-distributed-systems). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff or Principal Engineer, Distributed Systems - **Company:** Amazon.com, Inc. - **Location:** Seattle, WA, United States - **Salary:** $160,000.0 - $200,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Law Practice Management Software, System Configuration, Data Stores, Distributed Systems, NoSQL, Reliability Engineering, Software Engineering, Unstructured Data, Large Language Models, Low Latency, Deployment Automation, Production Code - **Published:** September 13, 2026 - **Apply:** https://www.careerjet.com/job/us31291c10dc13406e49307e615a1fa352/eaa ## About the Role Demonstrated experience building and operating high-volume, low-latency services on shared infrastructure. Hands-on distributed systems experience, including writing production code. Experience with multi-tenancy and workload isolation. Familiarity with admission control, queuing, scheduling, and backpressure mechanisms. Experience managing tail latency. Reliability engineering and failure recovery experience. Experience with large relational and/or NoSQL data stores. Experience owning production infrastructure and deployment automation. Nice to Have: Experience with AI/LLM pipelines or token-processing infrastructure. Track record of rearchitecting legacy or high-risk systems. Experience using AI tooling to accelerate implementation. Legal technology or document-processing domain familiarity. ## Description The Staff or Principal Engineer, Distributed Systems will own a major vertical of distributed infrastructure from design through production operations. The organization runs large-scale services executing millions of queries and processing hundreds of millions of tokens per minute across substantial structured and unstructured data volumes. A core challenge of this role is multi-tenant workload isolation, with customer workloads varying by three or more orders of magnitude in size and demand while sharing infrastructure. This is a hands-on production engineering role requiring direct ownership of system design, production code, infrastructure, deployment, observability, reliability, and incident response., Own a major vertical of distributed infrastructure from infrastructure definitions through production operations. Design and implement systems for workload isolation, admission control, fairness, and tail-latency management under multi-tenant contention. Write production code and directly investigate performance and reliability problems. Own infrastructure-as-code, deployment configuration, production promotion, and observability for the systems you build. Lead incident response for services within your ownership scope. Identify and remove bottlenecks across ingestion, storage, retrieval, orchestration, and AI-processing pipelines. Establish technical patterns and working implementations that raise the architectural quality of the broader engineering team., Are you a software development engineer who wants to build the systems that hundreds of thousands of AWS customers rely on to keep their cloud environments secure and compliant? Th… + 1 day ago ## Related Videos - [Fault Tolerance and Consistency at Scale: Harnessing the Power of Distributed SQL Databases](https://www.wearedevelopers.com/videos/1520-fault-tolerance-and-consistency-at-scale-harnessing-the-power-of-distributed-sql-databases) - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [Swapping Low Latency Data Storage Under High Load](https://www.wearedevelopers.com/videos/746-swapping-low-latency-data-storage-under-high-load) - [Reliable scalability: How Amazon.com scales on AWS](https://www.wearedevelopers.com/videos/983-reliable-scalability-how-amazon-com-scales-on-aws) - [Database Magic behind 40 Million operations/s](https://www.wearedevelopers.com/videos/748-database-magic-behind-40-million-operations-s) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)