Database Reliability Engineer

Cloudlinux
Spain
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
5 years minimum
Working hours
Shift work
Job source

Tech stack

Artificial Intelligence Airflow Amazon Web Services Business Logic Audit Trail Microsoft Azure Cloud Computing Computer Programming Databases Data Infrastructure Shard (Database Architecture) Linux
+28 more
DevOps Disaster Recovery Distributed Systems Federal Information Processing Standards (FIPS) Network Packet Python (Programming Language) PostgreSQL MongoDB Operational Databases Redis Ansible SQL Databases Apache Zookeeper Network Storage Google Cloud Gerrit Grafana HybridCloud Gitlab Kubernetes Bare Metal Data Analytics Apache Kafka Vertica Terraform Jenkins Nvme Golang

Job description

CloudLinux is transforming the Linux infrastructure market by ensuring security and stability for over 500,000 servers worldwide. Our products - CloudLinux OS, TuxCare, and Imunify360 - are the de facto standard in the hosting industry and Enterprise segment. We are seeking a visionary engineer to lead the evolution of our data platform. In 2025, we are shifting from classic database administration to an Internal Database-as-a-Service (DBaaS) model. We need a specialist who doesn’t just “configure backups,” but designs resilient distributed systems, writes code to automate infrastructure, and transforms databases into a reliable service for product teams. If you are tired of endless tickets and want to build platforms capable of processing petabytes of data, this role is for you. Your Challenges & Responsibilities - DBaaS Architecture: Design and implement a self-service platform based on Terraform and Ansible, enabling the deployment of HA clusters (PostgreSQL and ClickHouse, of Kubernetes operators for stateful workloads. - Expertise & Mentorship: Serve as the technical authority for product teams, helping them optimize data schemas and SQL queries for high-load systems. Our Tech Stack - Databases: PostgreSQL 15+ (Patroni, PgBouncer), ClickHouse (Sharded/Replicated), MongoDB, Redis, Kafka - Data & Analytics: Apache Airflow, Redash (Infrastructure & Integration) - Infrastructure: Own 3+ DC colocation (OpenNebula, Kubernetes, Bare Metal), AWS, Google Cloud, Azure, DO - Hybrid Cloud - Automation & IaC: Terraform, Ansible, Python/Go, GitLab, Jenkins, Gerrit - Observability: VictoriaMetrics, Grafana, Loki Why CloudLinux? - Culture: A Remote-first company with an “Employees First” principle. We value results, not hours in the office. - Impact: Your architectural decisions will determine the stability of services used by thousands of companies around the world. - Growth: We support professional development and pay for training and

Requirements

conferences. Requirements What We Expect From You - AI-Augmented Engineering: You don’t view AI as a replacement for deep technical fundamentals, but as a high-leverage tool. We actively use AI agents to automate boilerplate, analyze complex logs, and speed up research. We expect you to be open to modern workflows and integrate AI into your day-to-day operations, allowing you to focus your brainpower on the true architectural challenges. - Deep PostgreSQL Expertise (5+ years): You know MVCC internals, understand locking mechanics, can configure Patroni and PgBouncer with your eyes closed, and have experience with seamless major version upgrades under load. - ClickHouse Mastery: Experience operating large clusters, understanding ZooKeeper/ClickHouse Keeper, sharding, replication internals, and the ability to diagnose performance issues at the data-part level. - Engineering Mindset (SRE/DevOps): You hate doing the same task twice by hand. Experience writing complex Terraform modules and Ansible roles is mandatory. Programming skills in Python or Go for automation are a huge plus. - Hybrid Environment Experience: You understand the differences between running DBs on Bare Metal vs. Kubernetes vs. Cloud and know how to optimize TCO and disk subsystem performance (NVMe, Network Storage). - Systems Approach: You see the big picture - from the network packet to the application business logic. You understand the importance of security (FIPS, Audit logs) and Disaster Recovery. Nice to Have - Experience building an Internal Developer Platform (IDP). - Experience operating databases in Kubernetes (CloudNativePG, Altinity Opera

About the company

MongoDB, Redis) in a heterogeneous environment (Bare Metal + OpenNebula + Kubernetes + Public Clouds). You will turn infrastructure into a product. - Scaling ClickHouse: Manage exponentially growing analytics clusters (12+ clusters, tens of terabytes of data). You will tackle sharding, table engine optimization (ReplicatedMergeTree), and building reliable S3 backup pipelines under high load. - Data Platform & Analytics Support: Maintain and scale the infrastructure for Apache Airflow and Redash. You will ensure the reliability of ETL pipelines and visualization tools, bridging the gap between raw infrastructure and the data analytics team. - Reliability as Code: Implement SRE practices in data management. Replace manual incident response with automated self-healing mechanisms. Define and implement SLO/SLI for all databases. - Stack Modernization: Lead the migration process from legacy solutions to modern cloud patterns. Participate in decision-making regarding the implementation

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on es.trabajo.org

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard ¡ WWC 2025

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo ¡ LIVE

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade ¡ LIVE

3:42 min

Comparing in-memory and Redis storage for cache scalability

Simone Sanfratello ¡ WWC 2022

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 ¡ WWC Europe 2026

Videos

See all

Related articles

See all