SRE Manager / SRE Architect

Qode LLC
Fort Mill, SC, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Microsoft Azure Bash Shell Cloud Computing Cloud Engineering Continuous Delivery DevOps Distributed Systems Github Monitoring of Systems Python (Programming Language) OpenShift
+25 more
Windows PowerShell Release Management Reliability Engineering Site Reliability Engineering Practices Prometheus Software Deployment Datadog Scripting Cloud Platform System System Availability Grafana Infrastructure as Code (IaC) Gitlab-ci Kubernetes Infrastructure Automation Frameworks Information Technology Deployment Automation Performance Monitor Cloudwatch Terraform Splunk Dynatrace Docker Elk Stack Jenkins

Job description

We are seeking a highly experienced and hands-on Site Reliability Engineering (SRE) Manager / SRE Architect to lead reliability, availability, performance, and release management initiatives across enterprise-scale applications and platforms. This role requires a strong blend of SRE, DevOps, Release Management, Cloud Engineering, Automation, and Production Operations expertise.

The ideal candidate will be deeply involved in designing and implementing reliability strategies, driving release governance, improving deployment processes, and ensuring operational excellence across cloud-native environments., Site Reliability Engineering (SRE)

  • Design and implement SRE best practices focused on reliability, scalability, performance, and availability.
  • Define and monitor SLIs, SLOs, and error budgets across critical applications and services.
  • Drive proactive monitoring, alerting, observability, and incident management processes.
  • Lead root cause analysis (RCA) efforts and implement preventive measures.
  • Improve system resiliency through automation, self-healing capabilities, and operational excellence.
  • Establish reliability standards across distributed systems and cloud platforms.

Release Management

  • Own and drive end-to-end release management processes across multiple environments.
  • Coordinate application releases across development, QA, UAT, staging, and production environments.
  • Develop release governance, release calendars, deployment strategies, rollback procedures, and change management processes.
  • Partner with development, QA, infrastructure, and business teams to ensure smooth production deployments.
  • Identify and mitigate release risks while minimizing downtime and business impact.
  • Implement deployment automation and continuous delivery best practices.

DevOps & Automation

  • Design and maintain CI/CD pipelines using modern DevOps tools.
  • Automate infrastructure provisioning, deployment, monitoring, and operational workflows.
  • Drive Infrastructure as Code (IaC) adoption using Terraform or similar technologies.
  • Support cloud-native architectures and containerized application deployments.
  • Partner with engineering teams to improve developer productivity and deployment velocity.

Cloud & Platform Engineering

  • Manage and optimize cloud infrastructure on AWS and/or Azure.
  • Support Kubernetes, container orchestration, and cloud-native application platforms.
  • Ensure platform scalability, security, compliance, and operational readiness.
  • Drive platform modernization initiatives and operational transformation efforts.

Requirements

Do you have experience in Terraform?, Core SRE Skills

  • 15+ years of IT experience with strong focus on SRE, DevOps, Platform Engineering, or Production Support.
  • Extensive hands-on experience implementing SRE practices in enterprise environments.
  • Strong understanding of:
  • SLI/SLO/Error Budgets
  • Incident Management
  • Problem Management
  • Capacity Planning
  • Reliability Engineering
  • Observability & Monitoring

Release Management

  • Proven experience managing large-scale production releases.
  • Strong expertise in:
  • Release Planning
  • Release Governance
  • Change Management
  • Deployment Automation
  • Rollback Strategies
  • Production Readiness Reviews

DevOps & Cloud

  • Hands-on experience with:
  • AWS and/or Azure
  • Kubernetes (EKS, AKS, OpenShift preferred)
  • Docker
  • Terraform
  • GitHub Actions, Jenkins, Azure DevOps, GitLab CI/CD
  • Experience building and maintaining CI/CD pipelines.

Monitoring & Observability

  • Strong experience with:
  • Dynatrace
  • Datadog
  • Splunk
  • Prometheus
  • Grafana
  • ELK Stack
  • CloudWatch

Scripting & Automation

  • Experience with Python, Bash, PowerShell, or similar scripting languages.
  • Strong automation mindset with focus on operational efficiency.

Nice to Have

  • LaunchDarkly end-to-end implementation experience
  • Feature flag management and progressive delivery strategies.
  • Financial Services, Banking, or Wealth Management domain experience.
  • Experience leading SRE or DevOps transformation initiatives.
  • Cloud certifications (AWS, Azure, Kubernetes).

Preferred Candidate Profile

  • Strong hands-on SRE leader, not just a people manager.
  • Deep expertise in Release Management and Production Support.
  • Proven background in DevOps, Cloud Engineering, and Platform Reliability.
  • Ability to work with development, infrastructure, security, and business teams.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all