> Markdown version of [/jobs/ext/2070823-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2070823-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Arango Contextual Data Platform - **Location:** Madrid, Spain - **Contract:** Permanent contract - **Skills:** Clean Code Principles, Artificial Intelligence, Amazon Web Services, Software Applications, Bash Shell, Big Data, Cloud Computing, Cloud Engineering, Computer Programming, Continuous Integration, Data Infrastructure, DevOps, Disaster Recovery, Distributed Data Store, Distributed Systems, Fault Tolerance, Monitoring of Systems, Python (Programming Language), Linux Kernel, Reliability Engineering, Prometheus, Circleci, Data Logging, Scripting, Google Cloud, Cloud Platform System, System Availability, Grafana, Reliability of Systems, Git, Data Layers, Containerization, Kubernetes, Production Code, Terraform, Software Version Control, Docker, Elk Stack, Jenkins, Golang, Programming Languages - **Published:** August 15, 2026 - **Apply:** https://www.adzuna.es/contact-us.html ## About the Role * Proven experience as an SRE or DevOps Engineer in a cloud-native environment. * Proficiency with Kubernetes in managing large-scale, distributed systems. * Experience with cloud providers such as AWS and Google Cloud (GCP). * Solid understanding of networking, security practices, and troubleshooting methods * Understanding of Linux internals (processes, environment variables etc.) * Familiarity with containerization technologies (e.g., Docker). * Knowledge of CI/CD practices and tools (Jenkins, CircleCI, etc.). * Familiarity with alerting, monitoring and observability tools (e.g., Prometheus, Grafana, ELK stack). * Strong troubleshooting and problem-solving skills, with the ability to address complex infrastructure issues. * Excellent communication and collaboration skills with a focus on continuous improvement and operational excellence. * Strong ability to self-organize and to work independently as part of a remote team *Knowledge of version control systems, particularly Git * Familiarity with programming languages such as Golang or Python Nice-to-Have: * Experience managing distributed databases or large-scale data storage systems. *Knowledge of security best practices in cloud environments. * Experience with scripting languages like Python or Bash. * Experience with Infrastructure-as-Code (IaC) tools like Terraform is a plus. *Experience working with GitOps * Strong programming skills in Golang, with experience in developing automation tools, scripts, or services. Location: EU Timezone, preferably within the EU itself (Remote)- Spain, Portugal, etc. ## Description At ArangoDB, we are building a robust, cloud-native infrastructure to support our distributed database systems, which power mission-critical applications for a wide range of industries. We are searching for a Site Reliability Engineer (SRE) to ensure the reliability, scalability, and performance of our infrastructure and applications, with a focus on automation, monitoring, and optimizing cloud environments. As a Site Reliability Engineer (SRE), you will be responsible for maintaining and improving the reliability of our distributed database systems running on Kubernetes and cloud environments (AWS, Google Cloud). You will design, implement, and maintain scalable infrastructure solutions, improve and expand observability into these solutions, and troubleshoot complex system issues. It is expected that you will come to work to write clean and efficient code in Golang, working closely with development teams Your goal is to ensure high availability and performance of our cloud-based systems, automating repetitive tasks, and enhancing our CI/CD pipelines. If you're passionate about building resilient systems, managing cloud infrastructure, and using Golang to create scalable solutions (or willingness to learn Golang), we want to hear from you!, * Design, implement, and maintain cloud infrastructure on AWS and Google Cloud platforms. * Ensure the scalability, performance, and reliability of our Kubernetes-based distributed database systems. * Collaborate with developers to write efficient, production-grade code in Golang to automate infrastructure management and improve system operations. * Optimize and automate CI/CD pipelines, deployment processes, and monitoring systems to support our production environment. * Develop strategies for disaster recovery, high availability, and fault tolerance. * Proactively identify system bottlenecks, troubleshoot, and resolve issues across the stack (network, OS, cloud infrastructure). * Implement monitoring, logging, and alerting systems to ensure visibility into system health and performance. * Participate in on-call rotations to support critical production systems and respond to incidents. * Collaborate with cross-functional teams to improve overall system reliability and scalability. * Collaborate with the Customer Success team to resolve customer issues., At Arango, we believe that AI is only as powerful as the data foundation. Our mission is to help organizations build AI systems that can reason, decide and act based on unified, current, and trusted business context at scale. We are helping define a new category of infrastructure: the contextual data layer for AI. Working at Arango means: + Contributing to cutting-edge AI and data infrastructure + Collaborating with experienced engineers, marketers, and product leaders + Helping shape how enterprises build AI-powered applications If you're excited about the intersection of AI, data, and social media, we'd love to hear from you. Powered by JazzHR Inscribirse en esta oferta Recibir ofertas similares por correo electrónico Al crear una alerta, aceptas nuestros Términos y condiciones y Política de privacidad, y el uso de cookies. ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [7 Most Popular Web Developer Jobs in Europe](https://www.wearedevelopers.com/magazine/163-7-most-popular-web-developer-jobs-in-europe) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries)