Senior Site Reliability Engineer (SRE)

Capgemini
Atlanta, GA, United States
2 months ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$76,918.0 - $120,203.0
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Application Programming Interfaces (APIs) Application Performance Management App Store (IOS) Microsoft Azure Distributed Systems Node.Js Reliability Engineering Backend Performance Monitor Front End Software Development Splunk
+2 more
Api Management Microservices

Job description

node.js problem management Microsoft Azure DevOps & CI/CD java SDK Integration microservices architecture Latency Optimization splunk

  • The Senior Site Reliability Engineer (SRE) is a hands-on role responsible for the availability, performance, and end-to-end observability of QSR digital platforms across Mobile (iOS/Android), Web, and POS systems.
  • This role is part of the Observability team and works closely with mobile, web, and backend engineering teams to ensure full visibility into customer journeys and user experience.
  • The focus is on building and operating Real User Monitoring (RUM), synthetic monitoring, and end-to-end telemetry correlation using Splunk and SignalFx, ensuring issues are detected before customer impact. This is not a monitoring-only role-it requires active involvement in instrumentation, release observability, and reliability engineering., * Define and enforce SLIs, SLOs, and error budgets for critical customer journeys (ordering, checkout, payments)
  • Own end-to-end observability across Mobile, Web, and POS platforms
  • Implement and operate RUM and synthetic monitoring for customer-facing journeys
  • Build mobile-first monitoring coverage including app performance, crash rates, API performance, and user journey tracking
  • Use Splunk and SignalFx to design dashboards, detectors, and actionable alerts
  • Enable correlation across mobile CDN API backend systems using logs, metrics, and traces
  • Partner with engineering teams for instrumentation, SDK integration, and embedding observability into releases
  • Analyze telemetry to detect post-release issues, device/OS-specific failures, and network degradation
  • Lead response for P1/P2 incidents and drive root cause analysis
  • Automate operational toil and improve reliability, * Clear visibility into mobile, web, and POS customer journeys
  • Issues identified before customer complaints or app store feedback
  • Strong, low-noise user-impact-driven alerting
  • Reduced crash rates, latency, and checkout failures
  • Observability embedded into every release

Requirements

  • Strong experience with Splunk (logs, dashboards) and SignalFx (metrics/APM)
  • Hands-on with RUM and synthetic monitoring tools
  • Experience with mobile observability (iOS/Android), including performance monitoring and crash analysis
  • Strong understanding of distributed systems and microservices (Java, Node.js)
  • Experience with Azure (AKS, App Services, APIM)
  • Ability to correlate frontend issues with backend services
  • Experience with CI/CD pipelines and observability in release processes

Benefits & conditions

The pay range that the employer in good faith reasonably expects to pay for this position is $36.98/hour - $57.79/hour. Our benefits include medical, dental, vision and retirement benefits. Applications will be accepted on an ongoing basis.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:06 min

Developer experience and project variety at scale

Alexandra Petri · World Congress 2023

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

45 sec

Working securely with Node.js path application programming interfaces

Sonya Moisset · World Congress 2023

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:12 min

Choosing TypeScript for complex backend applications

Maximilian Otto Maximilian Otto · World Congress 2024

3:55 min

Identifying underlying Node.js runtime vulnerabilities using fuzzing tools

Sonya Moisset · World Congress 2023

Videos

See all

Related articles

See all