> Markdown version of [/jobs/ext/339765-platform-operations-manager-devops-site-reliability-engineering](https://www.wearedevelopers.com/jobs/ext/339765-platform-operations-manager-devops-site-reliability-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Platform Operations Manager (DevOps & Site Reliability Engineering) - **Company:** CareADHD - **Location:** London, UK - **Experience:** Expert - **Salary:** £75,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Cloud Computing, Cloud Engineering, Databases, Continuous Integration, DevOps, Disaster Recovery, Fault Tolerance, Github, Monitoring of Systems, PostgreSQL, Linux System Administration, Networking Basics, Node.Js, Performance Tuning, Release Management, Reliability Engineering, Prometheus, TypeScript, Datadog, AWS Cdk, Data Logging, Cloud Platform System, System Availability, Grafana, AWS Lambda, Event Driven Architecture, Gitlab-ci, Kubernetes, Infrastructure Automation Frameworks, Deployment Automation, Performance Monitor, Cloudwatch, Api Gateway, Terraform, Serverless Computing, Docker, Pagerduty, Jenkins, Microservices - **Published:** June 11, 2026 - **Apply:** https://uk.indeed.com/viewjob?jk=049ff045d98d531a ## About the Role Do you have experience in TypeScript?, * 8+ years of experience in DevOps, Platform Engineering, SRE, or Infrastructure Engineering * Proven experience leading operational or platform engineering teams * Strong experience managing distributed or offshore technical teams * Experience supporting business-critical production systems with high availability requirements * Experience operating cloud-native platforms in AWS environment Technical Skills Strong hands-on experience with: * AWS cloud infrastructure and services * CI/CD pipeline design and automation * Infrastructure as Code (Terraform or AWS CDK) * Kubernetes and container orchestration * Monitoring, logging, and observability platforms * Incident management and operational support * Linux systems administration and networking fundamentals Strong understanding of: * Site Reliability Engineering principles * High availability and disaster recovery design * Platform scalability and resilience * Security and operational governance * Performance optimisation and capacity planning Experience with tools such as: * Terraform * GitHub Actions / GitLab CI / Jenkins * CloudWatch * Datadog / Grafana / Prometheus * Docker / Kubernetes * PagerDuty or similar incident management tooling Leadership Competencies Operational Leadership Strong ownership mindset with the ability to lead operational stability and platform reliability across the organisation. Communication Excellent communication and stakeholder management skills, particularly across distributed engineering teams. Problem Solving Calm and effective under pressure with strong incident management and troubleshooting capabilities. Collaboration Works effectively across engineering, product, QA, and security teams to support reliable platform delivery. ## Description We believe that high-quality data and meaningful insight are essential to improving clinical services, understanding patient journeys, and ensuring that care is delivered efficiently and effectively., We are looking for an experienced and hands-on Platform Operations Lead to own the reliability, availability, performance, and operational stability of Care ADHD's technology platforms. This role combines DevOps, Site Reliability Engineering (SRE), cloud infrastructure, platform operations, and technical leadership - ensuring that our systems are securely deployed, highly available, scalable, and operational 24/7/365. You will lead platform operations across both the UK and India, working closely with engineering, QA, security, and product teams to ensure our infrastructure and deployment capabilities support a fast-moving and high-quality engineering organisation. This is a highly technical leadership role requiring someone who is equally comfortable defining operational strategy, improving engineering practices, and being hands-on with cloud infrastructure, automation, monitoring, incident response, and reliability engineering., Platform Reliability & Operations * Own the operational health, availability, and reliability of all production and non-production environments * Ensure platforms are monitored, maintained, and operational 24/7/365 * Lead platform incident management, root cause analysis, and service recovery processes * Establish and improve operational readiness, resilience, and disaster recovery capabilities * Define and manage SLAs, SLOs, and operational performance metrics * Ensure high levels of platform uptime, stability, scalability, and security DevOps & Infrastructure Engineering * Design, build, and maintain cloud infrastructure primarily within AWS * Lead infrastructure automation and Infrastructure as Code initiatives using Terraform or AWS CDK * Design and optimise CI/CD pipelines to support efficient, secure, and reliable software delivery * Improve deployment automation, release management, and environment consistency Support engineering teams with platform tooling, deployment strategies, and operational best practices * Drive improvements in: Deployment reliability Infrastructure scalability Platform security Cost optimisation Operational efficiency Site Reliability Engineering (SRE) * Implement and maintain observability solutions including: Monitoring Logging Alerting Tracing * Develop proactive approaches to incident prevention and operational resilience * Lead reliability engineering practices including: Capacity planning Performance monitoring Fault tolerance High availability design * Reduce operational toil through automation and self-service tooling * Establish strong incident response and post-incident review processes Leadership & Team Management * Lead and mentor platform operations and DevOps engineers across the UK and India * Build a collaborative, accountable, and high-performing operational culture * Allocate and coordinate operational resources across projects and platform priorities * Work closely with the Director of Engineering to align platform strategy with product and engineering delivery goals * Collaborate with engineering leads, QA, security, and product teams to support platform and release readiness Security, Compliance & Governance * Ensure infrastructure and operational processes follow security best practices * Support compliance with GDPR and healthcare-related operational standards * Help implement operational governance, access controls, and infrastructure security policies * Work closely with security and engineering teams to manage vulnerabilities and operational risk Technology Environment * AWS cloud infrastructure * Kubernetes and containerised services * Serverless platforms (AWS Lambda, API Gateway) * Node.js / TypeScript applications * PostgreSQL and cloud-native databases * Terraform / AWS CDK * CI/CD pipelines and deployment automation * Monitoring and observability tooling * Microservices and event-driven architectures ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Platform Engineering vs. DevOps Why not both?](https://www.wearedevelopers.com/videos/885-platform-engineering-vs-devops-why-not-both) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [We adopted DevOps and are Cloud-native, Now What?](https://www.wearedevelopers.com/videos/485-we-adopted-devops-and-are-cloud-native-now-what) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) ## Related Articles - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)