> Markdown version of [/jobs/ext/2667048-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2667048-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Plytix PIM - **Location:** Málaga, Spain (Remote available) - **Contract:** Temporary contract - **Skills:** Java (Programming Language), Amazon Web Services, Microsoft Azure, DevOps, Disaster Recovery, Monitoring of Systems, Network Security, Enterprise Messaging Systems, Reliability Engineering, Ansible, Prometheus, Mesos, Data Logging, Docker Swarm, Grafana, Backend, Kubernetes, Infrastructure Automation Frameworks, Deployment Automation, Apache Kafka, Terraform, Programming Languages - **Published:** September 2, 2026 - **Apply:** https://www.adzuna.es/contact-us.html ## About the Role We are looking for a skilled and experienced Site Reliability Engineer to join our team, and to help us scale and keep our applications reliable under certain parameters of quality (low latencies, low error rate, and so on). Our current technical stack is simple, so the ideal candidate should have a strong background in: + Grafana/Prometheus: Monitoring. + Elastic: Our logging backend., + 8+ years of experience in IT and 3+ years of experience in a Site Reliability Engineering or DevOps role. + Strong experience with AWS, Kubernetes, and logging/monitoring systems. + Experience with automation and configuration management tools (e.g. Ansible, Terraform). + Strong understanding of networking and security principles. + Excellent troubleshooting and problem-solving skills. + Ability to work independently and as part of a team in a fast-paced environment. + Strong communication and collaboration skills. Nice to have: + Experience with other cloud providers (e.g. GCP, Azure). + Experience with other container orchestration systems (e.g. Docker Swarm, Mesos). + Experience with other messaging systems (e.g. Kafka). + Experience with other programming languages (e.g. Go, Java). ## Description We're on the search for a Site Reliability Engineer (SRE) ready to join our SRE team in monitoring, automating, and alerting our systems. Our Development team is based in Malaga, but if you code your best from the comfort of your own home or even from a different city-the location doesn't matter, you do What is Plytix? This is where you'd typically see a long paragraph of boring text that no one reads, peppered with corporate buzzwords that don't mean anything. But Plytix is no typical company, so we'll spare you from that. Instead, watch this video, and if you like what you see, keep scrolling. https://www.youtube.com/watch?v=nGQDV5fNGp4, + Maintain reliable and scalable infrastructure on AWS using Kubernetes and other tools. + Monitor and troubleshoot production systems and respond to incidents in a timely manner. + Design and implement automated deployment and testing pipelines to ensure quality and reliability. + Perform capacity planning and scaling of infrastructure to support growing demand. + Collaborate with development teams to ensure that applications are designed for reliability and scalability. + Develop and maintain monitoring and alerting systems to detect issues before they become critical. + Continuously improve the reliability and performance of our infrastructure and applications. + Be part of releases. You'll need to collaborate to determine the suitability of the deployments. After 1 month: You'll have a good grasp of our favorite tools and processes, and your code will start flowing. You'll start learning our current architecture and services, performing small tasks so that you can start to show off your engineering skills. You'll take part in any engineering discussions and production procedures. After 3 months: You should know almost everything about our infrastructure and have an idea on what things should be improved. At this point you'll be ready to propose your ideas to improve our system: + Observability: Improve our logging system and make it easier for other teams to see the status of the platform. + Automation: Propose some automations for some processes. + High-availability: Prepare and test our system to make it more efficient and reliable. + Alerting: Improve our alerting system. + Disaster recovery: Have procedures for disaster recovery. 6 months in: You're fully integrated into the team at this point. You know all our strengths and weaknesses, and should be able to prepare a roadmap to improve all our systems. You're also collaborating with other departments to find the best possible solutions for our customers to have the best quality system. You're delivering high-quality, fast-running solutions that'll be tested to iron out any potential errors. By now, you're confident in your role, performing tasks in development as well as production. Who will you be working with? You'll be working with the Infrastructure team making sure that everything is running smoothly, but they aren't the only shining face you'll be working with on a daily basis. You'll also be working with the QA, Development, and DevOps teams to keep customers smiling about how amazing their Plytix PIM is. If you're curious about what it's like working at our office, take a peek into what Plytix has to offer (you know you want to ). ## Related Videos - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Nest.js - TypeScript in the backend can also be clean](https://www.wearedevelopers.com/videos/1033-nest-js-typescript-in-the-backend-can-also-be-clean) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)