Site Reliability Engineer (SRE)

Capgemini
Fort Mill, SC, United States
about 2 months ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Compensation
$76,918.0 - $120,203.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Application Layers Application Performance Management Confluence JIRA Microsoft Azure Cloud Engineering Software Documentation DevOps Monitoring of Systems Octopus Deploy Performance Tuning
+17 more
Reliability Engineering Prometheus UML Google Cloud System Availability Delivery Pipeline Grafana Mttr Multi-Cloud Infrastructure as Code (IaC) Performance Monitor Teamcity Terraform Splunk Dynatrace Atlassian Bamboo Jenkins

Job description

JIRA & Confluence aiops monitoring tools observability aws cloud Site Reliability Engineering (SRE) DevOps & CI/CD MTTR Google Cloud Platform (GCP) dynatrace, * Lead the design and implementation of full-stack observability solutions with Dynatrace as the primary platform.

  • Configure Dynatrace for application performance monitoring (APM), infrastructure monitoring, and intelligent alerting.
  • Build advanced dashboards and integrate Dynatrace with event management systems to enable proactive incident prevention and root cause analysis.
  • Collaborate with teams to optimize Dynatrace usage for AIOps-driven insights and automated anomaly detection.
  • Provide oversight for production operations to maximize reliability and automation.
  • Develop and evolve SRE best practices, runbooks, and tooling to ensure high availability and resilience.
  • Implement data-driven operational strategies to improve decision-making and reduce MTTR.
  • Hands-on experience with Dynatrace, Splunk, ELK, Grafana, Prometheus, and (future) ThousandEyes.
  • Build and manage CI/CD pipelines and Infrastructure as Code (IaC) solutions using Terraform, Jenkins, TeamCity, Octopus, Bamboo, and U-Deploy across hybrid/multi-cloud environments.
  • Develop and manage DevOps pipelines in AWS, Azure, and GCP using Terraform and cloud-native tooling.
  • Strong developer background with the ability to understand application layers and infrastructure interactions.
  • Define and document standard operating procedures, architecture diagrams, and system documentation using Jira, Confluence, and UML.
  • Identify areas for process and efficiency improvement within Platform Services Operations; recommend and implement solutions.
  • Drive automation initiatives across all operational processes.
  • Proactively monitor system capacity and health indicators; provide analytics and forecasts for scaling.

Requirements

We are looking for a Site Reliability Engineer with deep expertise in Dynatrace and a strong background in observability, automation, and cloud operations. This role focuses on designing and implementing highly reliable, scalable solutions while driving proactive monitoring and operational excellence., * Expert-level experience with Dynatrace, including dashboard creation, alert configuration, and integration with other observability tools.

  • Strong knowledge of AIOps, performance tuning, and proactive incident management.
  • Familiarity with hybrid/multi-cloud environments and modern DevOps practices.
  • Excellent problem-solving skills and ability to work in a fast-paced, collaborative environment.

Benefits & conditions

The pay range that the employer in good faith reasonably expects to pay for this position is $36.98/hour - $57.79/hour. Our benefits include medical, dental, vision and retirement benefits. Applications will be accepted on an ongoing basis.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · World Congress 2021

2:13 min

Evaluating UML, block diagrams, and the C4 model

Simon Lasselsberger · World Congress 2022

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

3:07 min

Establishing service level agreements directly for internal platforms

Pawel Piwosz · LIVE

4:36 min

Visualizing system architecture with Mermaid UML diagrams

Jakov Semenski · LIVE

Videos

See all

Related articles

See all