SRE
Infinity Quest
UK
2 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.adzuna.co.uk
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
£48,103.0
Working hours
Regular working hours
Job source
Tech stack
Artificial Intelligence
Amazon Web Services
Microsoft Azure
Distributed Systems
Information Technology Operations
Python (Programming Language)
Machine Learning
Reliability Engineering
Ansible
Datadog
Mttr
Containerization
+3 more
Dynatrace
Docker
Microservices
Job description
- Work closely with Product Engineering team and implement strategies for modernizing IT operations enhancing observability and toil reduction.
- Architect and deploy observability platforms to monitor system health, performance, and reliability effectively.
- Propose & drive strategies for AI-driven alerting and proactive anomaly detection to reduce MTTD & MTTR.
- Develop and enforce SRE best practices, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets.
- Establish & create AIOPS roadmap for improving operational efficiency.
- Lead efforts to automate repetitive tasks (toil) using scripting, orchestration tools, and AI/ML-based solutions.
- Drive toil automation initiatives for automated incident responses & self-healing automation for achieving autonomous operations.
- Collaborate with cross-functional teams to ensure systems are scalable, resilient, and maintainable.
- Drive incident management and root cause analysis processes through automation, ensuring continuous improvement to enable autonomous operations.
- Partner with engineering, architecture, and product teams to enable shift-left engineering practices ensuring reliability.
- Mentor and guide teams on adopting SRE principles and tools.
- Advocate for a culture of reliability, automation, and continuous improvement across the organization.
Requirements
- Strong expertise in implementing Site Reliability Engineering (SRE) principles.
- Advanced knowledge of establishing observability using tools Dynatrace & Datadog (primary skills).
- Proficiency in automation & scripting using Python & Ansible (primary skills).
- Strong experience with cloud platforms AWS & Azure (primary skills).
- Solid understanding of containerization and orchestration tools like Docker and Kubernetes.
-
Proficiency in cloud native distributed systems & microservices architecture.
- Exposure to AI/ML techniques for predictive analytics and automated problem resolution.
- Familiarity with CI/CD pipelines & enabling automated release & deployment engineering solutions.
- Good to have experience with chaos engineering tools like Gremlin or Chaos Monkey and implementing automation frameworks for resilience tracking.
- Ability to manage and prioritize multiple projects in a fast-paced environment.
- Strong interpersonal and communication skills to work effectively across teams.
- Excellent problem solving, analytical thinking, and adaptability.
Strategic mindset balancing engineering excellence with business priorities.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.adzuna.co.uk
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
AJ
Austin Joy
over 4 years ago
BR
Benjamin Ruschin
Navigating the AI Shift
about 1 year ago
LM
Luis Minvielle
Is Software Engineering Over-Saturated?
over 2 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
over 2 years ago
LM
Luis Minvielle
Why Upskilling And Reskilling is Important For Developers
over 2 years ago