> Markdown version of [/jobs/ext/2642718-site-reliability-engineer-sre](https://www.wearedevelopers.com/jobs/ext/2642718-site-reliability-engineer-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer (SRE) - **Company:** Stellent IT LLC - **Location:** Charlotte, NC, United States - **Experience:** Expert - **Contract:** Temporary contract - **Skills:** Microsoft Windows, Artificial Intelligence, Bash Shell, Cluster Analysis, Linux, Middleware, IBM WebSphere MQ, Python (Programming Language), Enterprise Messaging Systems, Windows PowerShell, Reliability Engineering, Site Reliability Engineering Practices, Prometheus, Software Vulnerability Management, Scripting, Cloud Platform System, Grafana, Event Driven Architecture, Kubernetes, Low Latency, Apache Kafka, Splunk, Dynatrace, Confluent - **Published:** August 21, 2026 - **Apply:** https://www.dice.com/job-detail/ece3fbf8-2959-474b-b8f4-1d0462b1eb03 ## About the Role Please share qualified candidates for the SRE Lead position with strong hands-on experience in IBM MQ and Kafka/Confluent, large-scale messaging/production engineering, SRE practices, monitoring/observability, incident management, RCA, SLI/SLO, HA/resiliency, and Shell/Python/PowerShell automation. Linux/Windows experience is required, while Kubernetes and Banking/Financial Services experience is preferred., * Observability tools (Dynatrace, Splunk, Prometheus, Grafana) for messaging platforms * Event and anomaly detection in high-volume systems * Strong scripting/automation skills: * Shell, Python, PowerShell * Experience managing Linux/Unix and Windows production environments Knowledge of: * Event-driven architecture and messaging-based integration patterns Understanding of: * Messaging platform security (TLS, certificates, channel auth, encryption) * Vulnerability remediation and risk mitigation in production systems * Excellent troubleshooting skills in high-pressure, real-time environments (e.g., message backlog, latency spikes, connection failures * Experience implementing SRE frameworks (SLIs, SLOs, error budgets) specifically for messaging workloads Familiarity with: * Kubernetes / containerized messaging platforms * Experience with: * Kafka ecosystem components (Schema Registry, Connect, Streams) * IBM MQ advanced features (Native HA, clustering) Exposure to: * AI-driven operations (AIOps), anomaly detection, or automated remediation * Large-scale messaging modernization or migration programs * Messaging or middleware certifications (IBM MQ, Kafka, or equivalent) * Experience in regulated environments (e.g., financial services) ## Description * Leading reliability engineering for high-scale messaging platforms supporting tens of thousands of runtimes and high-volume message throughput * Driving EOL remediation, patching, and stabilization across MQ queue managers and Kafka clusters Implementing SRE best practices: * SLIs / SLOs focused on message delivery, latency, and availability * Incident management, escalation, and postmortem culture * Enhancing observability and monitoring for messaging flows, queue depths, lag, and throughput * Designing proactive fault detection and auto-remediation strategies (e.g., DLQ handling, backlog mitigation, failover recovery) * Building resilient messaging platforms capable of supporting real-time, event-driven workloads * Supporting global production messaging environments with on-call rotation and escalation ownership * Partnering with engineering, application, and security teams tensure reliability, scalability, and secure message transport * Strong experience in Site Reliability Engineering / Production Engineering Hands-on expertise with: * IBM MQ (queue managers, clustering, channels, DLQ management) * Kafka / Confluent platform (topics, brokers, partitions, consumer groups) * Large-scale distributed messaging systems and runtime management ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)