> Markdown version of [/jobs/ext/1957109-cockroachdb-database-engineer-site-reliability-engineer-sre](https://www.wearedevelopers.com/jobs/ext/1957109-cockroachdb-database-engineer-site-reliability-engineer-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # CockroachDB Database Engineer / Site Reliability Engineer (SRE) - **Company:** Skysoft Inc - **Location:** Austin, TX, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Cloud Computing, Databases, Continuous Integration, Database Security, DevOps, Disaster Recovery, Distributed Data Store, Distributed Systems, Identity and Access Management, Python (Programming Language), PostgreSQL, Performance Tuning, Query Optimization, Reliability Engineering, Site Reliability Engineering Practices, Ansible, Prometheus, SQL Databases, Management of Software Versions, Datadog, Data Logging, Scripting, Google Cloud, Cloud Platform System, System Availability, Grafana, Database Optimization, Indexer, Database Migration, Containerization, Kubernetes, Infrastructure Automation Frameworks, Low Latency, Terraform, Splunk - **Published:** August 6, 2026 - **Apply:** https://www.dice.com/job-detail/ba0f06cf-bf18-4290-a0ce-e9912f19e64e ## About the Role Database Technologies Strong hands-on experience with CockroachDB Administration Expertise in distributed SQL databases and cluster management Database performance tuning and query optimization Backup, recovery, replication, and data protection strategies High Availability and Disaster Recovery architecture Site Reliability Engineering (SRE) Strong understanding of SRE principles and operational excellence Experience defining and tracking SLIs, SLOs, and Error Budgets Incident response, RCA, and reliability engineering practices Production monitoring, observability, and capacity management Reliability automation and operational process improvement Cloud & Automation Experience with AWS, Azure, or Google Cloud Platform Infrastructure as Code (Terraform, Ansible, etc.) Linux/Unix administration Scripting using Python, Shell, or Go CI/CD pipeline integration and automation Preferred Qualifications Experience supporting large-scale, mission-critical distributed systems. Knowledge of Kubernetes and containerized deployments. Experience with observability platforms such as Prometheus, Grafana, ELK, Datadog, or Splunk. Understanding of security, compliance, and governance requirements for database platforms. CockroachDB certification or equivalent distributed database expertise is highly desirable. Soft Skills Strong analytical and problem-solving capabilities. Excellent stakeholder communication and collaboration skills. Ability to work independently in a fast-paced production environment. Strong ownership mindset with a focus on reliability and customer experience. Nice-to-Have Experience with PostgreSQL internals (CockroachDB compatibility layer). Experience with distributed systems, consensus mechanisms, and multi-region architectures. Exposure to FinTech, Retail, E-Commerce, or large-scale cloud-native platforms. ## Description We are seeking a highly skilled CockroachDB Database Engineer with strong Site Reliability Engineering (SRE) experience to design, implement, manage, and optimize large-scale distributed database platforms. The ideal candidate will have hands-on expertise in CockroachDB administration, performance tuning, high availability, disaster recovery, automation, observability, and operational reliability. The role requires close collaboration with development, infrastructure, and platform engineering teams to ensure highly available, resilient, and scalable database services., Design, deploy, administer, and maintain production-grade CockroachDB clusters across cloud and on-premises environments. Monitor database health, performance, latency, throughput, and resource utilization to ensure service reliability and availability. Implement and manage backup, restore, disaster recovery, and business continuity strategies. Perform database capacity planning, performance tuning, indexing, and query optimization. Develop automation scripts and Infrastructure-as-Code (IaC) solutions to streamline provisioning, upgrades, and operational tasks. Establish and manage SRE practices including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets. Drive incident management, root cause analysis (RCA), postmortems, and preventive remediation activities. Build and maintain monitoring, logging, and alerting solutions using tools such as Prometheus, Grafana, ELK, Datadog, or similar platforms. Collaborate with DevOps and Engineering teams to improve platform reliability, scalability, security, and operational excellence. Support production releases, database migrations, version upgrades, and platform modernization initiatives. Participate in on-call rotation and provide support for critical production incidents. Implement database security controls, access governance, auditing, and compliance best practices. ## Related Videos - [How building an industry DBMS differs from building a research one](https://www.wearedevelopers.com/videos/768-how-building-an-industry-dbms-differs-from-building-a-research-one) - [Optimizing Discovery: PostgreSQL's Role in Transforming GetYourGuide's Search](https://www.wearedevelopers.com/videos/1647-optimizing-discovery-postgresql-s-role-in-transforming-getyourguide-s-search) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Crew Management System for Airlines: Plan duties for pilots & flight attendants worldwide](https://www.wearedevelopers.com/videos/1444-crew-management-system-for-airlines-plan-duties-for-pilots-flight-attendants-worldwide) - [Dynamic Entities in .NET: Building Low-Code Systems on Top of Entity Framework Core](https://www.wearedevelopers.com/videos/100218-dynamic-entities-in-net-building-low-code-systems-on-top-of-entity-framework-core) ## Related Articles - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers)