SRE - PERM REMOTE

Insight Global
Dunwoody, GA, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Working hours
Regular working hours
Job source

Tech stack

.NET Framework Application Performance Management Microsoft Azure Cloud Computing Cloud Engineering Program Optimization Information Systems DevOps Log Analysis Performance Tuning Reliability Engineering YAML
+8 more
Software Repository Software Troubleshooting Azure Powershell Infrastructure Automation Frameworks Information Technology Deployment Automation Bicep Software Version Control

Job description

We are seeking a Site Reliability Engineer (SRE) / DevOps Engineer to join a growing team responsible for supporting and scaling a Microsoft Azure environment. This role is ideal for an engineer who enjoys blending infrastructure, automation, deployment support, and reliability engineering to ensure highly available, scalable systems.

As the organization continues to grow its customer base, this individual will play a key role in improving observability, monitoring system health, diagnosing issues, and proactively preventing outages before they occur. While the position currently leans more heavily toward DevOps and infrastructure support, it will evolve into a balanced SRE/DevOps role with increased ownership of reliability initiatives, capacity planning, root cause analysis, and system performance optimization.

Responsibilities

Support and maintain cloud infrastructure within Microsoft Azure

Build and manage CI/CD pipelines and deployment automation

Monitor application and system health using Azure observability tools

Analyze logs, diagnose incidents, and troubleshoot production issues

Improve monitoring, alerting, and overall system observability

Partner with engineering teams to improve application reliability and performance

Implement Infrastructure-as-Code (IaC) solutions

Participate in preventative maintenance, system optimization, and capacity planning efforts

Contribute to a culture of proactive reliability engineering and operational excellence

We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global’s Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.

Requirements

2-5 years of experience in a DevOps Engineer, Site Reliability Engineer (SRE), Cloud Engineer, or related role

Experience working within Microsoft Azure environments

Strong troubleshooting, problem-solving, and systems-thinking abilities

Experience supporting applications in a .NET ecosystem

Azure & Cloud Technologies

Azure DevOps

Azure Monitor

Application Insights

Log Analytics

Azure CLI

Azure Container Apps

Infrastructure & Automation

YAML-based pipeline deployments

Bicep

Git repositories and source control best practices

Infrastructure-as-Code (IaC) experience

Reliability & Operations

Log analysis and troubleshooting

Incident diagnosis and resolution

System health monitoring

Strong observability mindset Kusto Query Language (KQL)

ARM Templates

Additional Azure platform expertise

Experience implementing or managing observability solutions

Capacity planning experience

Root cause analysis and post-incident review experience

Previous ownership of SRE initiatives or reliability programs

Bachelor’s degree in Computer Science, Engineering, Information Systems, or a related field

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:56 min

Provisioning a secure container infrastructure with Bicep

Matthias Falkenberg +1 · WWC 2022

1:38 min

Managing and versioning system prompts as YAML files

Kevin Lewis Kevin Lewis +1 · WWC 2025

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

4:09 min

Selecting infrastructure tools and determining proper abstraction layers

Alayshia Knighten Alayshia Knighten · WWC 2024

Videos

See all

Related articles

See all