> Markdown version of [/jobs/ext/1691400-site-reliability-developer](https://www.wearedevelopers.com/jobs/ext/1691400-site-reliability-developer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Developer - **Company:** Watchguard Sre - **Location:** Madrid, Spain - **Contract:** Permanent contract - **Skills:** Java (Programming Language), JIRA, Automation of Tests, Cloud Computing, Code Review, Computer Programming, Software Debugging, DevOps, Elasticsearch, Github, Python (Programming Language), Object-Oriented Software Development, Reliability Engineering, Software Construction, Software Engineering, Apache Spark, HybridCloud, Cloudformation, Kubernetes, Apache Flink, Build Process, Functional Programming, Software Coding, Terraform, New Relic (SaaS), Software Version Control, Serverless Computing, Docker, Jenkins, Artifactory, Microservices - **Published:** July 13, 2026 - **Apply:** https://www.jobleads.com/es/job/eede325e91113921925b79a3ce33d849f ## About the Role You are a customer-focused, data-driven developer who has a passion for delivering the best customer experience possible. You enjoy the thrill of coordinating and troubleshooting production issues and want to proactively find and fix issues. You have an understanding of cloud technologies, automation, everything-as-code, networking, microservice architectures, object-oriented design, SRE and DevOps cultures, proficiency in Python, Java, or Go programming and a desire to learn others. You come with proven knowledge of software engineering best practices for the full software development lifecycle including coding standards, code reviews, security, source control management, build processes, automated testing, deployment, monitoring, chaos engineering, and automated self-healing operations. As well as knowledge of tools and technologies like CloudFormation, Terraform, New Relic, Lambda, Serverless, Elasticsearch, Docker, Kubernetes, Spark, Flink, Jenkins, GitHub, Artifactory, Jira, etc. You are able to lead production incident response and postmortems through your strong analytical and problem-solving abilities as well as verbal and written communication skills. ## Description The WatchGuard SRE team owns the reliability and security of our production cloud environments alongside our application development teams to ensure we deliver the best possible experience to our customers. As you learn more about our systems, you will be: * Ensuring smooth production operations with development teams and leading large-scale event response. * Defining operational and security policies, standards, and processes for our development teams to follow. * Guiding our development teams through the process of establishing, monitoring, and achieving their service level agreements through the definition of service level indicators and objectives., * Working side-by-side with our application teams in production AWS, Azure, and hybrid cloud environments to ensure proper monitoring, security, reliability, automation, and support are in place. * Driving an operational excellence culture throughout WatchGuard with the simplification, automation, analysis, and evolution of our activities and processes. * Championing security and operational best practices to become known as a cloud expert by the rest of our development teams located across the globe. * Striving to provide the best possible customer experience even when things go wrong by participating in our on-call rotation and then coordinating and leading the production troubleshooting efforts. * Using your programming skills to develop automation or assist with debugging and fixing complex production issues. * Being curious, learning new things, and then sharing your knowledge through documentation, presentations, and guidance to other teams. ## Related Videos - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Collaboration Quantified: Lessons from Open Source Developer Networks](https://www.wearedevelopers.com/videos/1422-collaboration-quantified-lessons-from-open-source-developer-networks) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [The 8 Best Code Testing Tools](https://www.wearedevelopers.com/magazine/402-the-8-best-code-testing-tools) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Now is the time for industrialized software development](https://www.wearedevelopers.com/magazine/601-now-is-the-time-for-industrialized-software-development) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)