Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+5 more
Job description
We are looking for a Site Reliability Engineer to join a platform team responsible for operating and evolving a large-scale NoSQL database environment that supports business-critical services used by millions of users worldwide. This role sits at the intersection of Software Engineering, Site Reliability Engineering, and Platform Engineering. You will work on reliability, scalability, automation, and modernization of distributed database systems, while helping reduce operational toil through software development and infrastructure automation. The team owns the database platform end-to-end and is responsible for ensuring high availability, performance, capacity optimization, and operational excellence across both cloud and hybrid environments., * Operate, maintain, and optimize large-scale NoSQL database environments.
- Improve platform reliability through SLI/SLO-driven engineering practices.
- Develop automation and tooling to reduce operational workload and manual processes.
- Design and implement scalable solutions for distributed database platforms.
- Support database migrations, upgrades, and modernization initiatives.
- Contribute to capacity planning, performance tuning, and cost optimization.
- Build, maintain, and improve infrastructure using Infrastructure as Code principles.
- Participate in incident response, troubleshooting, root cause analysis, and postmortems.
- Collaborate closely with software engineers, platform teams, and stakeholders across the organization.
- Continuously evaluate emerging technologies and propose improvements to the platform landscape.
Requirements
- Strong experience with Cassandra and Dynamo
- Experience operating distributed database systems in production environments.
- Hands-on AWS experience.
- Strong Terraform knowledge.
- Programming experience in Python and Java
- Understanding of SLI, SLO, SLA, and reliability engineering principles.
- Experience with automation, observability, monitoring, and alerting.
- Knowledge of Elasticsearch or DynamoDB is highly beneficial.
- Strong communication skills and a collaborative mindset.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Find a Developer Job: 12 Best Job Sites For Developers
Dev Digest 120 - Apple and peers
Fully Remote Software Engineer Jobs
Is Software Engineering Over-Saturated?