> Markdown version of [/jobs/ext/2976469-sr-site-reliability-engineer-job-in-holmdel](https://www.wearedevelopers.com/jobs/ext/2976469-sr-site-reliability-engineer-job-in-holmdel). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. Site Reliability Engineer job in Holmdel - **Company:** CentralReach, LLC - **Location:** Holmdel, NJ, United States - **Experience:** Expert - **Salary:** $160,000.0 - $180,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), .NET Framework, Microsoft Windows, Artificial Intelligence, Amazon Web Services, Cloud Computing, Continuous Integration, Linux, Github, Python (Programming Language), Release Management, Reliability Engineering, Prometheus, Software Engineering, Systems Architecture, Datadog, Cloud Platform System, Grafana, Gitlab, Containerization, Kubernetes, Data Analytics, Splunk, Jenkins, Programming Languages - **Published:** September 18, 2026 - **Apply:** https://jobs.diversity.com/career/2494164/sr-site-reliability-engineer-new-jersey-nj-holmdel ## About the Role If you have a passion for the future, enjoy and thrive in an agile, fast-moving, ever-changing startup environment, welcome and take on technical challenges of all shapes and sizes, have excellent interpersonal skill and sense of humor and enjoy rolling up your sleeves and jumping in, then read on!, * Experience with monitoring, APM, and observability tools such as Splunk, Prometheus, Datadog, and OpenTelemetry. * Experience implementing observability strategies for logs, metrics, and traces. * Strong understanding of CI/CD practices and tools such as Jenkins, GitHub Actions, GitLab, Argo, and Kargo. * Strong understanding of major cloud providers, preferably AWS, and cloud-native infrastructure concepts. * Strong understanding of containerization technologies, including Kubernetes and Helm. * Experience with one or more programming languages, such as Java, Python, or Go, and familiarity with .NET application development. * Strong understanding of Linux, Windows, software development, systems, networking, and cloud concepts. * Experience using AI to improve productivity and amplify technical skills. ## Description As a Sr. SRE, you will work closely with the key stakeholders in Software Engineering to drive adoption of modern reliability practices like SLOs, error budget policies, actionable alerts, incident retrospectives, chaos testing, and end-to-end ownership., * Own production reliability, including availability, latency, performance, capacity planning, monitoring, emergency response, and uptime for production environments. * Define, maintain, and improve SLOs, SLIs, error budgets, actionable dashboards, and observability practices. * Analyze, troubleshoot, and resolve operational issues that affect service reliability and SLO performance. * Build and automate multi-environment observability capabilities, including capacity forecasting based on usage patterns. * Reduce toil and increase development velocity through automation and continuous improvement. * Provide production support, including incident, change, and problem management root cause analysis service restoration runbooks and standard operating procedures. * Identify data-driven opportunities to improve system architecture, availability, performance, and reliability. * Collaborate with software engineering teams on release management, roadmap planning, and operational readiness. * Implement and manage reliability and observability tools such as Datadog, Prometheus, and Grafana. ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [GitLab CI pipelines for a whole company](https://www.wearedevelopers.com/videos/143-gitlab-ci-pipelines-for-a-whole-company) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Enabling automated 1-click customer deployments with built-in quality and security](https://www.wearedevelopers.com/videos/83-enabling-automated-1-click-customer-deployments-with-built-in-quality-and-security) - [WeAreDevelopers LIVE - Modern DevOps for IoT Devices and More](https://www.wearedevelopers.com/videos/1805-wearedevelopers-live-modern-devops-for-iot-devices-and-more) ## Related Articles - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)