> Markdown version of [/jobs/ext/2412482-site-reliability-engineer-iii](https://www.wearedevelopers.com/jobs/ext/2412482-site-reliability-engineer-iii). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer III - **Company:** CME Group - **Location:** Belfast, UK - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Agile Methodology, Artificial Intelligence, Applications Architecture, Bash Shell, Cloud Computing, Cloud Engineering, Computer Programming, Continuous Integration, Data Distribution Service, Disaster Recovery, Distributed Systems, Domain Name System (DNS), Middleware, Monitoring of Systems, Hypertext Transfer Protocols (HTTP), Python (Programming Language), Linux System Administration, Network Protocols, Role-Based Access Control, Reliability Engineering, Ansible, Prometheus, Service Discovery, TCP/IP, Scripting, Google Cloud, File Transfer Protocol (FTP), Load Balancing, Grafana, Reliability of Systems, Generative AI, Infrastructure as Code (IaC), Containerization, Kubernetes, Infrastructure Automation Frameworks, Low Latency, Apache Kafka, Database Replication, Terraform, Splunk, Golang - **Published:** August 28, 2026 - **Apply:** https://uk.indeed.com/viewjob?jk=fe6d704b321cafe0 ## About the Role * Engineering & Scripting Discipline: Programming and scripting skills in high-level languages such as Python, Go, Java, or Bash to construct production-grade tooling. * Cloud Native & Systems Fundamentals: Proficiency with Linux-based systems, distributed systems, containerization (Kubernetes/GKE), and public cloud platforms (GCP/GCE). * Infrastructure as Code (IaC): Understanding of modern CI/CD patterns and IaC tools such as Terraform, Ansible, or Kubernetes Config Connector (KCC). * Networking & Protocols: Knowledge of core systems and networking concepts (TCP/IP, UDP, HTTP, DNS, load balancing, and messaging protocols). * AI & Agentic Engineering: Forward-thinking approach to automation, leveraging Generative AI and Agents (e.g., Gemini) to optimize platform operations. * Analytical Problem-Solving: Data-driven mindset to troubleshoot complex, non-linear system behaviors in a fast-paced, high-pressure trading ecosystem. * Communication & Adaptability: Strategic communication skills to translate technical requirements for cross-functional teams, coupled with an eagerness to learn independently and collaboratively. Preferred Qualifications / Desirable * Observability Stack: Hands-on experience with telemetry tools such as OpenTelemetry, Splunk, Prometheus, and Grafana. * Agile Integration: Comfort working within Agile frameworks and collaborative software development lifecycles. * Certifications: GCP Professional Cloud Architect, Certified Kubernetes Administrator (CKA), or Certified Kubernetes Application Developer (CKAD). * Domain Expertise: Any experience in Financial Markets or other highly regulated, ultra-low latency, high-concurrency environments would be highly beneficial although not essential," ## Description The Role: CME Group is seeking a Site Reliability Engineer (SRE) III to engineer reliability for our Google Cloud (GCP) infrastructure, Middleware Platform Engineering team, and core technology foundations powering our Clearing, Risk, and derivatives applications. In this role, you will help build resilient, automated systems that combine ultra-low latency with high-concurrency performance, enabling CME's product teams to innovate safely at scale. You will work alongside senior engineers, mentor junior colleagues, engage in the dynamic operation of production systems, and assist in driving our cloud transformation., * Middleware & Application Architecture: Architect, operate, and support the migration of application platforms-including Messaging (Kafka, RedPanda, MQ, Pub/Sub), Service Discovery (Consul, Vault), and Data Distribution (SFTP/JScape)-to Google Cloud Platform. Manage cluster lifecycles, data replication, RBAC, and workload placement. * Observability & Monitoring Fabric: Design, scale, and maintain our observability backbone using tools like OpenTelemetry, Splunk, Prometheus, and Grafana. Establish and continuously improve metrics, logs, alerting strategies, SLIs, and SLOs to enable fast issue detection. * Incident Response & Operations: Engage with urgency in live production incidents, take ownership of minor incidents, lead post-mortems, and ensure rapid system recovery. * Toil Reduction & Automation: Actively identify operational toil and eliminate manual effort through code, automation, and systematic platform improvements. * Resiliency & Testing: Contribute to disaster recovery (DR) strategies, continuous systems resiliency testing, and present reliability improvement suggestions to the Product backlog. * Collaboration & Leadership: Lead technical discussions for assigned scope, present solution options, collaborate across functional teams, and mentor junior SRE colleagues. ## Related Videos - [Inside Bitpanda's Tech Stack: Scaling a European Fintech Leader - Markus Dorner](https://www.wearedevelopers.com/videos/1979-inside-bitpanda-s-tech-stack-scaling-a-european-fintech-leader-markus-dorner) - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)