Site Reliability Engineer

Insight Global
Atlanta, GA, United States
11 days ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$114,400.0 - $124,800.0
Working hours
Regular working hours
Job source

Tech stack

Microsoft Azure Bash Shell Cloud Computing DevOps Domain Name System (DNS) Hypertext Transfer Protocols (HTTP) Virtual Private Networks (VPN) Python (Programming Language) Microsoft SQL Server Windows PowerShell RabbitMQ Reliability Engineering
+13 more
Zero Trust Network Access TCP/IP Transport Layer Security Google Cloud Load Balancing Computer Network Technologies Firewalls (Computer Science) Containerization Kubernetes Cloudflare Terraform Splunk Appdynamics

Job description

Our client is looking for an SRE that will Lead the reliability, scalability, security, and operational excellence of customer-facing platforms across Azure, GCP, and Kubernetes environments.

The ideal candidate will be responsible for driving production stability through automation, observability, incident management, and continuous improvement initiatives.

Key Responsibilities

  • Lead platform reliability, availability, and performance initiatives.

  • Design and support cloud infrastructure in Azure and GCP.

  • Manage and optimize Kubernetes environments and containerized applications.

  • Implement observability and monitoring using Splunk, AppDynamics, and cloud-native tools.

  • Support Cloudflare, Zscaler, SQL Server, RabbitMQ, and enterprise networking components.

  • Lead major incident response, RCA, and problem management activities.

  • Develop automation and self-healing solutions to improve operational efficiency.

  • Collaborate with Engineering, Product, Security, and Infrastructure teams to enhance customer experience and platform stability.

  • Serve as a technical escalation point for critical production and customer issues.

Requirements

  • 5+ years of experience in SRE, DevOps, Cloud Operations, or Infrastructure Engineering.

  • Strong expertise in Azure, GCP, Kubernetes, Cloudflare, Splunk, AppDynamics, SQL Server, RabbitMQ, and Zscaler.

  • Solid networking knowledge (DNS, TCP/IP, HTTP/S, CDN, WAF, Load Balancing, SSL/TLS, Firewalls, VPNs).

  • Experience with automation and scripting (Python, PowerShell, Bash, Terraform).

  • Strong customer-facing communication and stakeholder management skills.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

5:02 min

Mapping distributed compute paradigms to modern vehicles

Joachim Werner · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

3:50 min

Queues in TCP stacks and continuous network connections

Clemens Vasters Clemens Vasters · World Congress 2022

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all