Site Reliability Engineer SRE (contract)

Wells Fargo
Charlotte, NC, United States
11 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Systems Engineering Build Automation Microsoft Azure Bash Shell DevOps Monitoring of Systems Python (Programming Language) Windows PowerShell Reliability Engineering Cloud Services Prometheus
+9 more
Scripting Google Cloud System Availability Grafana Containerization Splunk Appdynamics Dynatrace Servicenow

Job description

o Lead or participate in triaging and resolving alerts related to high availability applications and platforms o Serve was a key escalation point for complex production issues and coordinate recovery across teams o Develop and enhance monitoring dashboards mentions reporting and alert runbooks/playbooks o Analyze trends from historical incidents and help drive root cause analysis and long-term fixes o Build automation scripts and tools to improve response time and reduce manual intervention o Coordinate with application owners, Engineering teams and service providers to ensure timely and accurate resolution of issues o Participate in on-call support rotation as needed and contribute to building resilient operations model o Continuously improve onboarding documentation alert definitions and service catalogs to ensure support readiness

Requirements

In this contingent resource assignment, you may: Consult on or participate in moderately complex initiatives and deliverables within Systems Operations Engineering and contribute to large-scale planning related to Systems Operations Engineering deliverables. Review and analyze moderately complex Systems Operations Engineering challenges that require an in-depth evaluation of variable factors. Contribute to the resolution of moderately complex issues and consult with others to meet Systems Operations Engineering deliverables while leveraging solid understanding of the function, policies, procedures, and compliance requirements. Collaborate with client personnel in Systems Operations Engineering. Required Qualifications: Systems Engineering or Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work or consulting experience, training, military experience, education., * Applicants must be authorized to work for ANY employer in the U.S. This position is not eligible for visa sponsorship.

o Systems Engineering, Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education o Experience in handling incident response, troubleshooting live issues and restoring services quickly o Hands on experience with monitoring tools such as AppDynamics, CRIBL, Dynatrace, Splunk, Thousandeyes, Prometheus, Grafana etc o Strong experience working with ServiceNow or other ITSM platforms for Incident, Change and Problem management o Proven experience in writing runbooks playbooks and working across teams to streamline support processes o Familiarity with scripting languages such as PowerShell Python or bash for automation, o Prior experience in DevOps/SRE Environments or Platform teams supporting enterprise scale applications. Familiarity with container platforms and cloud services (Azure, AWS, GCP) o Experience with CI/CD pipelines deployment processes and observability best practices o Excellent communication and collaboration skills, with proactive approach to issue resolution Strong organizational and documentation skills to support knowledge-based creation and escalation paths

Benefits & conditions

Pulled from the full job description

  • Health insurance
  • Life insurance

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · WWC 2025

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all