CockroachDB Database Engineer / Site Reliability Engineer (SRE)

Skysoft Inc
Austin, TX, United States
7 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Microsoft Azure Cloud Computing Databases Continuous Integration Database Security DevOps Disaster Recovery Distributed Data Store Distributed Systems Identity and Access Management Python (Programming Language)
+25 more
PostgreSQL Performance Tuning Query Optimization Reliability Engineering Site Reliability Engineering Practices Ansible Prometheus SQL Databases Management of Software Versions Datadog Data Logging Scripting Google Cloud Cloud Platform System System Availability Grafana Database Optimization Indexer Database Migration Containerization Kubernetes Infrastructure Automation Frameworks Low Latency Terraform Splunk

Job description

We are seeking a highly skilled CockroachDB Database Engineer with strong Site Reliability Engineering (SRE) experience to design, implement, manage, and optimize large-scale distributed database platforms. The ideal candidate will have hands-on expertise in CockroachDB administration, performance tuning, high availability, disaster recovery, automation, observability, and operational reliability. The role requires close collaboration with development, infrastructure, and platform engineering teams to ensure highly available, resilient, and scalable database services., Design, deploy, administer, and maintain production-grade CockroachDB clusters across cloud and on-premises environments.

Monitor database health, performance, latency, throughput, and resource utilization to ensure service reliability and availability.

Implement and manage backup, restore, disaster recovery, and business continuity strategies.

Perform database capacity planning, performance tuning, indexing, and query optimization.

Develop automation scripts and Infrastructure-as-Code (IaC) solutions to streamline provisioning, upgrades, and operational tasks.

Establish and manage SRE practices including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.

Drive incident management, root cause analysis (RCA), postmortems, and preventive remediation activities.

Build and maintain monitoring, logging, and alerting solutions using tools such as Prometheus, Grafana, ELK, Datadog, or similar platforms.

Collaborate with DevOps and Engineering teams to improve platform reliability, scalability, security, and operational excellence.

Support production releases, database migrations, version upgrades, and platform modernization initiatives.

Participate in on-call rotation and provide support for critical production incidents.

Implement database security controls, access governance, auditing, and compliance best practices.

Requirements

Database Technologies

Strong hands-on experience with CockroachDB Administration

Expertise in distributed SQL databases and cluster management

Database performance tuning and query optimization

Backup, recovery, replication, and data protection strategies

High Availability and Disaster Recovery architecture

Site Reliability Engineering (SRE)

Strong understanding of SRE principles and operational excellence

Experience defining and tracking SLIs, SLOs, and Error Budgets

Incident response, RCA, and reliability engineering practices

Production monitoring, observability, and capacity management

Reliability automation and operational process improvement

Cloud & Automation

Experience with AWS, Azure, or Google Cloud Platform

Infrastructure as Code (Terraform, Ansible, etc.)

Linux/Unix administration

Scripting using Python, Shell, or Go

CI/CD pipeline integration and automation

Preferred Qualifications

Experience supporting large-scale, mission-critical distributed systems.

Knowledge of Kubernetes and containerized deployments.

Experience with observability platforms such as Prometheus, Grafana, ELK, Datadog, or Splunk.

Understanding of security, compliance, and governance requirements for database platforms.

CockroachDB certification or equivalent distributed database expertise is highly desirable.

Soft Skills

Strong analytical and problem-solving capabilities.

Excellent stakeholder communication and collaboration skills.

Ability to work independently in a fast-paced production environment.

Strong ownership mindset with a focus on reliability and customer experience.

Nice-to-Have

Experience with PostgreSQL internals (CockroachDB compatibility layer).

Experience with distributed systems, consensus mechanisms, and multi-region architectures.

Exposure to FinTech, Retail, E-Commerce, or large-scale cloud-native platforms.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:10 min

Synchronizing stateful application databases across multiple cloud providers

Alex Soto Alex Soto · WWC 2022

2:36 min

Analyzing limitations with PostgreSQL bitmap heap scans

Dharin Shah Dharin Shah · WWC 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:32 min

Generating functional runtime database columns using indexer properties

Halil İbrahim Kalkan Halil İbrahim Kalkan · WWC Europe 2026

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all