Site Reliability Engineer, Tech Services, Monetization Tech - USDS (Multiple Positions)

Amazon.com, Inc.
Bellevue, WA, United States
16 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Compensation
$129,960.0 - $246,240.0
Working hours
Regular working hours

Tech stack

Amazon Web Services Data as a Services Data Centers Fault Tolerance Performance Tuning Reliability Engineering Software Engineering Information Technology

Job description

Provide site reliability engineering support to ensure highest level of availability of large-scale, globally distributed, fault-tolerant ads systems. Engage in and improve the whole lifecycle of Ads systems, from system design consulting through launch reviews, deployment, operation and refinement. Build availability of services deployed across multiple data centers globally. Deliver tools/software to improve the reliability, scalability and operability of services, including designing, developing and deploying automation to sustainably scale with quality. Measure and monitor availability, latency and overall service health. Practice sustainable incident response and postmortems, performing root cause analysis of incidents to influence future product design and response activities., The future of software development is autonomous. AWS is building Kiro - a frontier developer tool that independently tackles development tasks, maintains context across interactio…

  • 7 hours ago

Requirements

Must have a Bachelor’s degree or foreign equivalent degree in Computer Science, Engineering (any), Information Technology, or a related field, and 2 years of related work experience. Of the required experience, must have 2 years of experience in each of the following: Providing functionality and reliability support for critical site components by measuring and monitoring availability, latency, and overall system health, including through performance tuning and troubleshooting; Monitoring system activity and resolving system issues; Coordinating and monitoring data services operations, including SLA management and system deployment; Analyzing error logs to identify issues and working with service owners to resolve issues, document their origins and develop future prevention mechanisms; and Creating and maintaining clear runbook instructions for services to use for alerts, troubleshooting and resolution.

About the company

TikTok USDS Joint Venture LLC is dedicated to the safety and security of millions of Americans who create, discover, and connect with what they love on the apps we operate. The Joint Venture has been established in compliance with the Executive Order signed by President Trump on September 25, 2025. Our foundation is a comprehensive data privacy and cybersecurity program. We operate under defined safeguards to protect national security and secure U.S. user data, apps and the algorithm. We safeguard the U.S. content ecosystem, holding decision-making authority for trust and safety policies and moderation. USDS Joint Venture helps ensure Americans can continue to express their creativity, discover new hobbies and interests, and build thriving communities and businesses on a global scale. Why Join Us Inspiring creativity is at the core of TikTok’s mission. Our innovative product is built to help people authentically express themselves, discover and connect - and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and bring joy - a mission we work towards every day. We strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. Every challenge is an opportunity to learn and innovate as one team. We’re resilient and embrace challenges as they come. By constantly iterating and fostering an “Always Day 1” mindset, we achieve meaningful breakthroughs for ourselves, our company, and our users. When we create and grow together, the possibilities are limitless. Join us. About the Team Our team plays a crucial role in ensuring the company’s success. We seek people who are willing to learn and put in the effort to solve problems. Our challenges are not your regular day-to-day problems - you’ll be part of a team that’s developing new solutions to new challenges. It’s working fast, at scale, and we’re making a difference. We are looking for talents to join us on this exciting journey!

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

51 sec

Repurposing hardware and operating underwater data centers

Chris Heilmann Chris Heilmann +1 · LIVE

1:05 min

Practical Byzantine Fault Tolerance in distributed computing systems

Jonan Scheffler · World Congress 2022

4:45 min

Validating subjective user reports with throwaway data services

Nico Seyboth Nico Seyboth +1 · World Congress 2025

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

4:03 min

Managing massive power consumption scaling in AI data centers

Stephan Gillich Stephan Gillich +3 · World Congress 2024

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

Videos

See all

Related articles

See all