Site Reliability Engineer (SRE)

Stellent IT LLC
Charlotte, NC, United States
15 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Microsoft Windows Artificial Intelligence Bash Shell Cluster Analysis Linux Middleware IBM WebSphere MQ Python (Programming Language) Enterprise Messaging Systems Windows PowerShell Reliability Engineering Site Reliability Engineering Practices
+12 more
Prometheus Software Vulnerability Management Scripting Cloud Platform System Grafana Event Driven Architecture Kubernetes Low Latency Apache Kafka Splunk Dynatrace Confluent

Job description

  • Leading reliability engineering for high-scale messaging platforms supporting tens of thousands of runtimes and high-volume message throughput
  • Driving EOL remediation, patching, and stabilization across MQ queue managers and Kafka clusters

Implementing SRE best practices:

  • SLIs / SLOs focused on message delivery, latency, and availability
  • Incident management, escalation, and postmortem culture
  • Enhancing observability and monitoring for messaging flows, queue depths, lag, and throughput
  • Designing proactive fault detection and auto-remediation strategies (e.g., DLQ handling, backlog mitigation, failover recovery)
  • Building resilient messaging platforms capable of supporting real-time, event-driven workloads
  • Supporting global production messaging environments with on-call rotation and escalation ownership
  • Partnering with engineering, application, and security teams tensure reliability, scalability, and secure message transport
  • Strong experience in Site Reliability Engineering / Production Engineering

Hands-on expertise with:

  • IBM MQ (queue managers, clustering, channels, DLQ management)
  • Kafka / Confluent platform (topics, brokers, partitions, consumer groups)
  • Large-scale distributed messaging systems and runtime management

Requirements

Please share qualified candidates for the SRE Lead position with strong hands-on experience in IBM MQ and Kafka/Confluent, large-scale messaging/production engineering, SRE practices, monitoring/observability, incident management, RCA, SLI/SLO, HA/resiliency, and Shell/Python/PowerShell automation. Linux/Windows experience is required, while Kubernetes and Banking/Financial Services experience is preferred., * Observability tools (Dynatrace, Splunk, Prometheus, Grafana) for messaging platforms

  • Event and anomaly detection in high-volume systems
  • Strong scripting/automation skills:
  • Shell, Python, PowerShell
  • Experience managing Linux/Unix and Windows production environments

Knowledge of:

  • Event-driven architecture and messaging-based integration patterns

Understanding of:

  • Messaging platform security (TLS, certificates, channel auth, encryption)
  • Vulnerability remediation and risk mitigation in production systems
  • Excellent troubleshooting skills in high-pressure, real-time environments (e.g., message backlog, latency spikes, connection failures
  • Experience implementing SRE frameworks (SLIs, SLOs, error budgets) specifically for messaging workloads

Familiarity with:

  • Kubernetes / containerized messaging platforms
  • Experience with:
  • Kafka ecosystem components (Schema Registry, Connect, Streams)
  • IBM MQ advanced features (Native HA, clustering)

Exposure to:

  • AI-driven operations (AIOps), anomaly detection, or automated remediation
  • Large-scale messaging modernization or migration programs
  • Messaging or middleware certifications (IBM MQ, Kafka, or equivalent)
  • Experience in regulated environments (e.g., financial services)

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all