> Markdown version of [/jobs/ext/636828-remote-site-reliability-engineer-sre](https://www.wearedevelopers.com/jobs/ext/636828-remote-site-reliability-engineer-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Remote Site Reliability Engineer (SRE) - **Company:** Air Ltd - **Location:** Cambridge, UK - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Systems Engineering, Microsoft Azure, Bash Shell, Cloud Computing, Cloud Engineering, Software Debugging, DevOps, Disaster Recovery, Distributed Systems, Failover, Fault Tolerance, Monitoring of Systems, Python (Programming Language), Linux System Administration, Networking Basics, Performance Tuning, Reliability Engineering, Prometheus, Software Engineering, Datadog, Data Logging, Pulumi, Scripting, Google Cloud, Load Balancing, System Availability, Grafana, Cloudformation, Containerization, Kubernetes, Infrastructure Automation Frameworks, Terraform, New Relic (SaaS), Docker - **Published:** June 25, 2026 - **Apply:** https://find.jobs/jobs-near-me/remote-site-reliability-engineer-sre-cambridge-cambridgeshire/2840381743-2/ ## About the Role * Around 4+ years of experience in Site Reliability Engineering (SRE), DevOps, or System Engineering. * Strong knowledge of cloud platforms (AWS, Azure, or GCP) and cloud-native architectures. * Experience with observability and monitoring tools (Prometheus, Grafana, ELK, Datadog, New Relic). * Proficiency in Infrastructure as Code (IaC) tools such as Terraform, CloudFormation, or Pulumi. * Hands-on experience with containerization and orchestration (Docker, Kubernetes, Helm). * Strong Linux system administration and networking fundamentals. * Experience with incident management, debugging, and root cause analysis. * Proficiency in scripting (Bash, Python, or Go) for automation and system monitoring. * Knowledge of load balancing, failover strategies, and distributed systems. * Understanding of security best practices, access control, and compliance requirements. * Strong communication skills and the ability to collaborate with cross-functional teams. ## Description As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the reliability, availability, and scalability of our systems. You will work at the intersection of software development and operations, implementing automation, monitoring, and performance optimization strategies to minimize downtime and improve system resilience. * This is a fully onsite position, based at our office in Lisbon, where you will collaborate closely with cross-functional teams in person and contribute to a dynamic and fast-paced environment. We are open to support with relocation efforts. Responsibilities * Design and implement scalable, reliable, and fault-tolerant systems across cloud environments. * Develop and maintain observability tools, including monitoring, logging, and alerting (e.g., Prometheus, Grafana, Datadog, ELK). * Automate infrastructure provisioning, deployment, and incident response using Infrastructure as Code (IaC) tools like Terraform or CloudFormation. * Optimize system performance, scalability, and incident response workflows to improve uptime. * Work closely with development and DevOps teams to improve system design for reliability. * Conduct root cause analysis (RCA) and implement preventative measures to minimize failures. * Ensure high availability by designing and maintaining load balancing, failover, and disaster recovery strategies. * Improve CI/CD pipelines to enhance deployment speed while maintaining stability. * Optimize cloud cost and resource utilization for AWS, Azure, or Google Cloud Platform (GCP). * Participate in on-call rotations to quickly address system failures and minimize downtime. ## Related Videos - [Unleashing Potential Across Teams: The Power of Infrastructure as Code](https://www.wearedevelopers.com/videos/930-unleashing-potential-across-teams-the-power-of-infrastructure-as-code) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Software Engineer Salary London](https://www.wearedevelopers.com/magazine/252-software-engineer-salary-london)