> Markdown version of [/jobs/ext/3042795-sre](https://www.wearedevelopers.com/jobs/ext/3042795-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SRE - **Company:** Infinity Quest - **Location:** UK - **Salary:** £48,103.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Distributed Systems, Information Technology Operations, Python (Programming Language), Machine Learning, Reliability Engineering, Ansible, Datadog, Mttr, Containerization, Dynatrace, Docker, Microservices - **Published:** September 24, 2026 - **Apply:** https://www.adzuna.co.uk/jobs/details/5895535242 ## About the Role * Strong expertise in implementing Site Reliability Engineering (SRE) principles. * Advanced knowledge of establishing observability using tools Dynatrace & Datadog (primary skills). * Proficiency in automation & scripting using Python & Ansible (primary skills). * Strong experience with cloud platforms AWS & Azure (primary skills). * Solid understanding of containerization and orchestration tools like Docker and Kubernetes. * Proficiency in cloud native distributed systems & microservices architecture. * Exposure to AI/ML techniques for predictive analytics and automated problem resolution. * Familiarity with CI/CD pipelines & enabling automated release & deployment engineering solutions. * Good to have experience with chaos engineering tools like Gremlin or Chaos Monkey and implementing automation frameworks for resilience tracking. * Ability to manage and prioritize multiple projects in a fast-paced environment. * Strong interpersonal and communication skills to work effectively across teams. * Excellent problem solving, analytical thinking, and adaptability. Strategic mindset balancing engineering excellence with business priorities. ## Description * Work closely with Product Engineering team and implement strategies for modernizing IT operations enhancing observability and toil reduction. * Architect and deploy observability platforms to monitor system health, performance, and reliability effectively. * Propose & drive strategies for AI-driven alerting and proactive anomaly detection to reduce MTTD & MTTR. * Develop and enforce SRE best practices, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets. * Establish & create AIOPS roadmap for improving operational efficiency. * Lead efforts to automate repetitive tasks (toil) using scripting, orchestration tools, and AI/ML-based solutions. * Drive toil automation initiatives for automated incident responses & self-healing automation for achieving autonomous operations. * Collaborate with cross-functional teams to ensure systems are scalable, resilient, and maintainable. * Drive incident management and root cause analysis processes through automation, ensuring continuous improvement to enable autonomous operations. * Partner with engineering, architecture, and product teams to enable shift-left engineering practices ensuring reliability. * Mentor and guide teams on adopting SRE principles and tools. * Advocate for a culture of reliability, automation, and continuous improvement across the organization. ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)