> Markdown version of [/jobs/ext/2729521-site-reliability-engineer-sre](https://www.wearedevelopers.com/jobs/ext/2729521-site-reliability-engineer-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer (SRE) - **Company:** (capital Markets) - **Location:** UK (Remote available) - **Contract:** Permanent contract - **Skills:** .NET Framework, Application Programming Interfaces (APIs), Artificial Intelligence, JIRA, Automation of Tests, Microsoft Azure, Bash Shell, Configuration Management, Code Coverage, Code Review, Relational Databases, Elasticsearch, Github, Python (Programming Language), PostgreSQL, Linux System Administration, Load Testing, Performance Tuning, Redis, Reliability Engineering, Prometheus, Software Requirements Analysis, TypeScript, Workflow Management Systems, Datadog, Data Logging, ReactJS, Istio, Grafana, Mttr, Caching, Reliability of Systems, Indexer, Backend, Containerization, Kubernetes, Infrastructure Automation Frameworks, Performance Monitor, Graphql, Terraform, Software Version Control, Docker - **Published:** September 5, 2026 - **Apply:** https://startup.jobs/site-reliability-engineer-capital-markets-gateway-8001370 ## About the Role * Proven experience as a Site Reliability Engineer or similar role. * Proficiency in logging, metrics, and tracing frameworks (DataDog, Loki, Prometheus, OpenTelemetry). * Experience with cloud platforms (Azure preferred) and infrastructure-as-code tools (e.g., Terraform). * Strong programming and scripting skills (Python, Bash). * Proficiency in containerization technologies and orchestration tools (Docker, Kubernetes). * Understandingof Linux-based systems, networking, and security principles related to containerized applications. * Strong problem-solving and troubleshooting skills, with a passion for identifying and resolving complex technical issues. * Excellent communication and collaboration abilities. * Ability to thrive in a fast-paced, constantly evolving environment. * Experience with PostgreSQL monitoring and optimization (Optional/Nice to have). If you're passionate about building resilient financial systems, optimizing observability at scale, and solving real-world reliability challenges in capital markets, we'd love to have you on our team! ## Description CMG is looking for a Site Reliability Engineer (SRE) with a strong focus on monitoring, observability, and alerting to ensure the reliability, performance, and scalability of our infrastructure and applications. You will be responsible for designing, implementing, and maintaining monitoring solutions to provide visibility into system health and performance, proactively detect anomalies, and reduce incident response time. Our Engineering Team The CMG engineering team consists of domain experts who work collaboratively within a culture of cross-domain knowledge sharing. We value engineers who are passionate about modern technologies and best practices. Our engineers are encouraged to challenge the status quo and are constantly seeking improvement and efficiency in our code-base and platform. CMG engineers are empowered to explore solutions using bleeding edge technologies such as AI and bring recommendations to the table. We are in a period of making impactful engineering decisions. As part of our process, we believe in taking the time for research and prototyping - this is critical in making the right decisions. Given the experience of our team, we have naturally adopted best practices from local development, through code review and into production rollouts. Besides the standard pull requests, test automation, code coverage tracking, containerization, and one-click deployments we are constantly reviewing these foundational components to develop new best practices., * Design, implement, and maintain monitoring and observability solutions using tools like Prometheus, Grafana Stack (Loki/Grafana/Tempo/Alert Manager), Datadog, and OpenTelemetry. * Define and implement SLOs, SLIs, and error budgets to measure system reliability. * Develop and optimize dashboards, alerts, and reports for system performance and business metrics. Alerting & Incident Management * Design actionable alerting strategies to minimize noise and improve MTTR. * Integrate alerting systems with Jira. * Establish and refine runbooks for on-call teams to handle alerts efficiently. * Empower teams to ensure observability coverage and incident response practices. Performance Optimization * Analyze system performance metrics, identify bottlenecks, and implement optimizations to improve system efficiency, scalability, and cost-effectiveness. * Help conduct load testing and capacity planning to ensure systems can handle peak traffic loads. Automation and Tooling * Identify opportunities for automation and develop tools to streamline operational processes, such as fail-over, configuration management, and monitoring. * Implement monitoring and alerting systems within automations to detect and resolve issues proactively. Collaboration and Communication * Collaborate closely with cross-functional teams, including software engineers, operations, and infrastructure teams, to understand system requirements, provide technical guidance, and drive solutions. * Communicate effectively to stakeholders about system changes, incidents, and improvements. * Foment and spread SRE principles and practices across company., * Azure as an infrastructure provider. We are reviewing secondary cloud options. * Docker + Kubernetes for microservice orchestration using Istio service mesh. * PostgreSQL for relational db, ElasticSearch for indexing, Redis for caching. * DataDog, Grafana and OpenTelemetry for observability. * GitHub for our Version Control and CI (with our own runners). * CD: Harness and FluxCD. * Terraform and Terragrunt as IaaC. * Python and bash for scripting infrastructure. * React - We're all in on React - we maintain multiple single-page React apps. * TypeScript - 99% of our codebase is TypeScript. * Latest .NET version for our backend services. * GraphQL - Our standard for API communication is GraphQL served by our DotNet Back-End. ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [The 12 Best Jobs for Software Engineers](https://www.wearedevelopers.com/magazine/401-the-12-best-jobs-for-software-engineers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)