Sr. Site Reliability Engineer job in Holmdel

CentralReach, LLC
Holmdel, NJ, United States
16 days ago
Apply on jobs.diversity.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$160,000.0 - $180,000.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) .NET Framework Microsoft Windows Artificial Intelligence Amazon Web Services Cloud Computing Continuous Integration Linux Github Python (Programming Language) Release Management Reliability Engineering
+13 more
Prometheus Software Engineering Systems Architecture Datadog Cloud Platform System Grafana Gitlab Containerization Kubernetes Data Analytics Splunk Jenkins Programming Languages

Job description

As a Sr. SRE, you will work closely with the key stakeholders in Software Engineering to drive adoption of modern reliability practices like SLOs, error budget policies, actionable alerts, incident retrospectives, chaos testing, and end-to-end ownership., * Own production reliability, including availability, latency, performance, capacity planning, monitoring, emergency response, and uptime for production environments.

  • Define, maintain, and improve SLOs, SLIs, error budgets, actionable dashboards, and observability practices.

  • Analyze, troubleshoot, and resolve operational issues that affect service reliability and SLO performance.

  • Build and automate multi-environment observability capabilities, including capacity forecasting based on usage patterns.

  • Reduce toil and increase development velocity through automation and continuous improvement.

  • Provide production support, including incident, change, and problem management root cause analysis service restoration runbooks and standard operating procedures.

  • Identify data-driven opportunities to improve system architecture, availability, performance, and reliability.

  • Collaborate with software engineering teams on release management, roadmap planning, and operational readiness.

  • Implement and manage reliability and observability tools such as Datadog, Prometheus, and Grafana.

Requirements

If you have a passion for the future, enjoy and thrive in an agile, fast-moving, ever-changing startup environment, welcome and take on technical challenges of all shapes and sizes, have excellent interpersonal skill and sense of humor and enjoy rolling up your sleeves and jumping in, then read on!, * Experience with monitoring, APM, and observability tools such as Splunk, Prometheus, Datadog, and OpenTelemetry.

  • Experience implementing observability strategies for logs, metrics, and traces.

  • Strong understanding of CI/CD practices and tools such as Jenkins, GitHub Actions, GitLab, Argo, and Kargo.

  • Strong understanding of major cloud providers, preferably AWS, and cloud-native infrastructure concepts.

  • Strong understanding of containerization technologies, including Kubernetes and Helm.

  • Experience with one or more programming languages, such as Java, Python, or Go, and familiarity with .NET application development.

  • Strong understanding of Linux, Windows, software development, systems, networking, and cloud concepts.

  • Experience using AI to improve productivity and amplify technical skills.

Benefits & conditions

$160,000-$180,000 USD

Backed by Roper Technologies, Inc. (Nasdaq: ROP), CentralReach is entering an exciting phase of growth, innovation, and scale.

Recognized as one of the best places to work over 10 times by organizations such as Inc, Built In, and NJBIZ, our culture is centered around impact, inclusion, and flexibility. As a hybrid company with collaborative offices in Ft. Lauderdale, FL Holmdel, NJ and Verona, Italy, we foster a workplace where top talent can thrive and make a real difference in the lives of those we serve.

We offer competitive compensation, comprehensive health benefits, generous PTO, 401(k) matching, and paid parental leave to our full-time employees. Our team members also enjoy hybrid work schedules, career development support, wellness programs, and opportunities to give back through CR Cares&trade, our community engagement initiative.

Be part of a market leader driving the future of care. Explore opportunities at centralreach.com/careers.

About the company

CentralReach is a leading provider of autism and IDD care software for Applied Behavior Analysis (ABA), multidisciplinary therapy, and special education. Trusted by more than 200,000 users, we enable therapy providers, educators, and employers to scale the way they deliver ABA and related therapies with innovative technology, market-leading industry expertise, and world-class customer satisfaction.

The Platform Engineering group at CentralReach builds the underlying technologies that power our Public and Private Cloud Platforms worldwide. The group is responsible for storage, data infrastructure, IT, observability systems, DevOps, SRE, provisioning, compute, orchestration platform, internal tools, internal platforms (laptops, networks, systems etc.) and services - all the components that make up the CentralReach Platform.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.diversity.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:46 min

Introduction to the speaker and engineering background

Llywelyn Griffith-Swain · World Congress 2023

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

3:14 min

Testing and environment management in GitLab CI

Martin Beránek · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

2:30 min

Feature requests for future GitLab CI versions

Martin Beránek · LIVE

Videos

See all

Related articles

See all