Site Reliability Engineer

Key2Source INC
Charlotte, NC, United States
about 2 months ago
Apply on dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Amazon Web Services Application Integration Architecture Microsoft Azure Cloud Computing Databases Continuous Integration DevOps Distributed Systems Middleware Monitoring of Systems Python (Programming Language)
+14 more
Windows PowerShell Reliability Engineering Prometheus Scripting Google Cloud Enterprise Software Applications System Availability Grafana Reliability of Systems Gitlab-ci Splunk Appdynamics Dynatrace Jenkins

Job description

We are looking for an experienced Site Reliability Engineer (SRE) with strong Application Support expertise to support and enhance the reliability, stability, and performance of enterprise applications and platforms. The ideal candidate will bridge the gap between application support and reliability engineering by driving operational excellence, automation, and system resilience., * Provide L2/L3 application support for enterprise applications in production and non-production environments.

  • Monitor application health, system availability, and performance using observability and monitoring tools.
  • Troubleshoot and resolve application, middleware, and infrastructure-related issues within SLA timelines.
  • Collaborate with Development, DevOps, Cloud, and Infrastructure teams for deployments, releases, and platform improvements.
  • Perform root cause analysis (RCA) and implement permanent fixes for recurring application issues.
  • Automate operational tasks, monitoring, and incident response processes.
  • Support CI/CD pipelines and deployment activities across environments.
  • Maintain operational documentation, runbooks, and support procedures.
  • Participate in incident management, change management, and on-call support rotations.
  • Continuously improve system reliability, scalability, and operational efficiency.

Requirements

  • 5+ years of experience in Site Reliability Engineering and Application Support roles.
  • Strong experience supporting enterprise applications in distributed environments.
  • Hands-on experience with Linux/Unix systems and troubleshooting.
  • Experience with monitoring and observability tools such as Splunk, Dynatrace, Grafana, AppDynamics, ELK, or Prometheus.
  • Knowledge of cloud platforms such as AWS, Azure, or Google Cloud Platform.
  • Experience with scripting/automation using Python, Shell, or PowerShell.
  • Familiarity with CI/CD tools such as Jenkins, GitLab CI/CD, or Azure DevOps.
  • Understanding of APIs, middleware, databases, and application integration troubleshooting.
  • Strong analytical, problem-solving, and communication skills.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:02 min

Applying an ETL methodology to infrastructure configuration management

Axel Barbier · World Congress 2023

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

57 sec

Extracting API schemas automatically during continuous integration builds

Axel Barbier · World Congress 2023

1:06 min

Developer experience and project variety at scale

Alexandra Petri · World Congress 2023

Videos

See all

Related articles

See all