Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+3 more
Job description
This is an exciting opportunity for a Senior Site Reliability Engineer to play a key role in the reliability, scalability, and performance of a growing AI-enabled platform. You will work at the intersection of infrastructure, automation, and AI operations, helping to build resilient systems that support innovative products and services at scale. The Company They are a technology-led organisation focused on delivering high-quality digital experiences to a global customer base. With a strong emphasis on innovation, automation, and operational excellence, they continue to invest in modern cloud infrastructure and AI-driven capabilities. Their collaborative environment encourages knowledge sharing, continuous improvement, and technical ownership. This role offers the chance to influence platform reliability while supporting the adoption of emerging technologies. The Role
- Design, build, and maintain reliable, scalable cloud infrastructure.
- Improve platform availability, performance, and operational resilience.
- Develop and enhance monitoring, alerting, and observability capabilities.
- Automate infrastructure management and operational processes.
- Support the operation and reliability of AI and machine learning platforms.
- Collaborate with engineering teams to improve deployment, testing, and release processes.
- Drive incident response, root cause analysis, and continuous service improvements.
- Contribute to platform security, compliance, and best practice engineering standards.
Requirements
- Strong commercial experience in Site Reliability Engineering, Platform Engineering, or DevOps.
- Expertise with cloud platforms such as AWS, Azure, or Google Cloud.
- Strong knowledge of Infrastructure as Code tools, including Terraform or similar technologies.
- Experience with containerisation and orchestration technologies such as Docker and Kubernetes.
- Hands-on experience with monitoring, logging, and observability solutions.
- Strong understanding of CI/CD pipelines and automation practices.
- Experience supporting production environments with high availability requirements.
- Knowledge of AI, machine learning infrastructure, or AI platform operations would be highly beneficial.
- Excellent problem-solving, communication, and stakeholder engagement skills.
Benefits & conditions
- Competitive salary package.
- Performance-related incentives and comprehensive benefits.
- Hybrid working and flexibility to support work-life balance.
- Exposure to modern cloud technologies and AI platforms.
- Career progression opportunities within a growing technology function.
- A collaborative environment focused on learning, innovation, and continuous improvement.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Where To Find Software Engineering Jobs
Find a Developer Job: 12 Best Job Sites For Developers
Fully Remote Software Engineer Jobs
Is Software Engineering Over-Saturated?