Site Reliability Engineer (SRE)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+12 more
Job description
- Leading reliability engineering for high-scale messaging platforms supporting tens of thousands of runtimes and high-volume message throughput
- Driving EOL remediation, patching, and stabilization across MQ queue managers and Kafka clusters
Implementing SRE best practices:
- SLIs / SLOs focused on message delivery, latency, and availability
- Incident management, escalation, and postmortem culture
- Enhancing observability and monitoring for messaging flows, queue depths, lag, and throughput
- Designing proactive fault detection and auto-remediation strategies (e.g., DLQ handling, backlog mitigation, failover recovery)
- Building resilient messaging platforms capable of supporting real-time, event-driven workloads
- Supporting global production messaging environments with on-call rotation and escalation ownership
- Partnering with engineering, application, and security teams tensure reliability, scalability, and secure message transport
- Strong experience in Site Reliability Engineering / Production Engineering
Hands-on expertise with:
- IBM MQ (queue managers, clustering, channels, DLQ management)
- Kafka / Confluent platform (topics, brokers, partitions, consumer groups)
- Large-scale distributed messaging systems and runtime management
Requirements
Please share qualified candidates for the SRE Lead position with strong hands-on experience in IBM MQ and Kafka/Confluent, large-scale messaging/production engineering, SRE practices, monitoring/observability, incident management, RCA, SLI/SLO, HA/resiliency, and Shell/Python/PowerShell automation. Linux/Windows experience is required, while Kubernetes and Banking/Financial Services experience is preferred., * Observability tools (Dynatrace, Splunk, Prometheus, Grafana) for messaging platforms
- Event and anomaly detection in high-volume systems
- Strong scripting/automation skills:
- Shell, Python, PowerShell
- Experience managing Linux/Unix and Windows production environments
Knowledge of:
- Event-driven architecture and messaging-based integration patterns
Understanding of:
- Messaging platform security (TLS, certificates, channel auth, encryption)
- Vulnerability remediation and risk mitigation in production systems
- Excellent troubleshooting skills in high-pressure, real-time environments (e.g., message backlog, latency spikes, connection failures
- Experience implementing SRE frameworks (SLIs, SLOs, error budgets) specifically for messaging workloads
Familiarity with:
- Kubernetes / containerized messaging platforms
- Experience with:
- Kafka ecosystem components (Schema Registry, Connect, Streams)
- IBM MQ advanced features (Native HA, clustering)
Exposure to:
- AI-driven operations (AIOps), anomaly detection, or automated remediation
- Large-scale messaging modernization or migration programs
- Messaging or middleware certifications (IBM MQ, Kafka, or equivalent)
- Experience in regulated environments (e.g., financial services)
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Fully Remote Software Engineer Jobs
Find a Developer Job: 12 Best Job Sites For Developers
Where To Find Software Engineering Jobs