> Markdown version of [/jobs/ext/2934391-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2934391-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** Cribl, Inc. - **Location:** Salem, OR, United States (Remote available) - **Experience:** Expert - **Salary:** $141,800.0 - $195,000.0 - **Contract:** Internship / Graduate position - **Skills:** JavaScript (Programming Language), Amazon Web Services, Microsoft Azure, Cloud Computing, Cloud Engineering, Configuration Management, Continuous Delivery, Linux, Reliability Engineering, Cloud Services, Ansible, Prometheus, TypeScript, Cloud Platform System, Grafana, Software Security, Sentry, Data Management, Cloudwatch, Kibana, Software Coding, Terraform, Splunk, New Relic (SaaS), Pagerduty - **Published:** September 16, 2026 - **Apply:** https://www.jofdav.com/jobs/59721184-senior-site-reliability-engineer ## About the Role If reliability is your passion, and you have always had strong opinions on how to make things better and have the desire to build consensus around ideas, then let's talk!, * Proven experience designing, implementing, and operating observability systems for complex cloud-based platforms, with deep knowledge of best practices and a strong drive to implement them leveraging Cribl products. * Experience with Configuration Management and Infrastructure as a Code Tools like Terraform (preferred) or Ansible. Experience working with Cloud SDKs is also a plus. * Knowledge of cloud platforms (prefer AWS and Azure) and container + orchestration technologies. * Experience with APM and Observability and related tools such as, New Relic, Splunk, CloudWatch, Prometheus, Grafana/Kibana, Sentry etc. * Extensive experience with enterprise scale continuous delivery environments. * Development with JavaScript/Node.js/TypeScript in a Linux/Mac environment. * Experience with sustainable incident response in a blameless environment. * Background in Linux Systems Engineering. * Experience with Incident response related tools for instance, PagerDuty, FireHydrant, Blameless etc. * Comfortable with a high level of autonomy and working with a distributed team. * Knowledge of Cloud and application security best practices. * Strong knowledge of cloud design patterns for scale, data management, resiliency, etc. * A love for high quality and a knack for testing. * Opinions about business metrics, and SLOs. #LI-GV1 #LI-Remote ## Description * Engage with teams and improve service delivery and reliability across their entire lifecycle * Measure and monitor all production systems with an eye towards availability, latency and overall system health * Seek out the cause of errors and instability in our production cloud services and drive teams towards better operational excellence * Engage with product and platform teams to improve and evolve systems by lobbying for changes that improve reliability, resilience, and observability * Help identify and drive down toil with creative innovation and automation * This position will require stand-by, on-call, or off-hours duties ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Debug a Kubernetes Operator](https://www.wearedevelopers.com/videos/487-debug-a-kubernetes-operator) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Add Location-based Searching to Site with ElasticSearch](https://www.wearedevelopers.com/videos/77-add-location-based-searching-to-site-with-elasticsearch) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 131 - AI'm not sure about OSS](https://www.wearedevelopers.com/magazine/472-dev-digest-131-ai-m-not-sure-about-oss)