Cloud Operations Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+8 more
Job description
A core engineer who owns the day-to-day delivery, resilience and improvement of our online products and services. The Cloud Operations Engineer resolves incidents independently, builds and maintains automation and Infrastructure as Code, and helps move the team from reactive to proactive operations - keeping services secure, performant and reliable for our customers. Engineers predominantly specialise in one platform - either Microsoft Azure or Oracle Cloud Infrastructure (OCI) - bringing real depth in their platform while collaborating across both.
This role may include working as part of an on-call rota to provide out-of-hours support. On-call is a compensated rota, with the details set out in our on-call policy., * Own the day-to-day operation, availability, resilience, and continuous improvement of cloud-based products and services, ensuring they remain secure, compliant, reliable, and high performing.
- Monitor, diagnose, and resolve incidents across compute, storage, network, and identity platforms, ensuring timely restoration of service and minimal business impact.
- Manage and maintain cloud infrastructure using Infrastructure as Code (IaC) methodologies and tools, ensuring consistency, scalability, and operational efficiency.
- Design, build, and support CI/CD pipelines to enable efficient, automated, and reliable deployment processes.
- Ensure services operate within defined Service Level Objectives (SLOs) and error budgets, using monitoring and observability data to drive operational decisions.
- Perform backup, recovery, disaster recovery, and failover activities, including regular testing to meet established RTO and RPO requirements.
- Implement and maintain security controls, identity and access management standards, patch management processes, and vulnerability remediation activities.
- Manage secrets, credentials, and privileged access controls in line with security and governance requirements.
- Ensure compliance with organisational and regulatory standards, including ISO 27001, SOC 2, GDPR, and related governance requirements.
- Identify recurring issues, raise problem records, and implement preventative measures to improve service stability and reduce incident recurrence.
- Develop and implement automation solutions that reduce manual effort, improve operational efficiency, and enhance service reliability.
- Apply FinOps principles to optimise cloud resource utilisation and control operational costs.
- Maintain accurate runbooks, operational procedures, and technical documentation to support consistent service operations and knowledge sharing.
- Participate in on-call support rotations and respond to operational alerts and critical incidents as required.
- Collaborate with Operations, Governance, Security, and Engineering teams to support service delivery, security compliance, and continuous improvement initiatives.
- Mentor and support junior engineers, sharing technical expertise and promoting engineering best practices.
- Contribute to a proactive, collaborative, and continuous improvement culture focused on operational excellence and customer outcomes.
- Support capacity planning, performance optimisation, cost management, and other operational initiatives aligned with business and technology objectives.
Requirements
- Solid hands-on experience operating a cloud platform in production, with specialist depth in either Microsoft Azure or Oracle Cloud Infrastructure (OCI)
- Demonstrable incident resolution and troubleshooting across compute, storage, network and identity
- Working knowledge of Infrastructure as Code (e.g. Terraform or Bicep) and CI/CD pipelines
- Scripting proficiency (PowerShell, Bash or similar) for automation
- Hands-on with identity and access (e.g. Microsoft Entra / IAM), patching, vulnerability remediation and safe secrets management
- Experience working to service levels (SLAs/SLOs) and within compliance controls (e.g. ISO 27001, SOC 2, GDPR)
- Willingness to take part in a compensated on-call rota (see on-call policy), * An associate-level cloud certification (e.g. Azure Administrator AZ-104 or OCI Architect Associate)
- Experience with observability and monitoring tools, and with backup / disaster-recovery testing
- ITIL awareness
- Klipboard is embracing AI at pace across our products and ways of working. We’re looking for people who are curious about how AI can enhance productivity, decision-making and customer outcomes, and who are open to learning and adapting as this space evolves.
About the company
“At Klipboard we’ve introduced a flexible hybrid work policy, where employees spend three days in the office and two days working from home. This approach promotes a balanced work environment that combines office collaboration with the comfort and convenience of remote work.”
Klipboard provides specialist software, services and support to deliver fully integrated trading and business management solutions to companies in the distributive trade - wherever they are in the world. With a unique depth of knowledge and experience in ERP/SaaS solutions, Klipboard has a wide range of clients includes wholesalers, distributors, merchants and retailers from small traders to multinational enterprises. Klipboard has offices in the UK, Ireland, The Netherlands, South Africa, Kenya and North America. Our mission is simple: to design and deliver high performance, integrated ERP solutions that enable our distributive trade customers to source effectively, stock efficiently, sell profitably and service competitively.
Klipboard is a global, growing business that embraces AI and emerging technologies to enhance customer outcomes, collaboration, and continuous improvement. We’re looking for people who are curious about or fluid with AI, open to change, and excited to learn how technology can improve the way we work and help our customers which is always supported by strong human insight and communication., You may also have seen from our recent posts that we are excited to begin sharing our new company name - Klipboard. Kerridge Commercial Systems (KCS) is becoming Klipboard and our new brand is designed to bring together our expertise across distribution, automotive, retail, rental, transport management, manufacturing, and field service management. We have offices based across the world and we are looking for talented individuals to join our growing teams. Due to our growth over the last few years it is an exciting time to join us as we enter our next chapter! At Klipboard we’ve introduced a flexible hybrid work policy, where employees spend three days in the office and two days working from home. This approach promotes a balanced work environment that combines office collaboration with the comfort and convenience of remote work.”
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
7 Cloud Computing Trends Coming in 2025 for Developers
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production
Navigating the AI Shift
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud