Site Reliability Engineer Manager ( Healthcare Domain)

Vaarida Technologies Llc
New York, NY, United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Microsoft Azure Continuous Integration DevOps Site Reliability Engineering Practices Prometheus Google Cloud Cloud Platform System Grafana Mttr Splunk Appdynamics

Requirements

10+ years in engineering, operations, or SRE roles

5+ years leading SRE, platform, or reliability-focused teams

Proven experience implementing SRE practices at scale (SLIs, SLOs, error budgets)

Strong background in cloud environments (AWS, Azure, Google Cloud Platform)

Hands-on experience with observability tools (Splunk, AppDynamics, Prometheus, etc.)

Experience in incident management and production operations at scale

Ability to operate effectively in high-pressure and complex enterprise environments

Preferred Qualifications

Experience driving organizational transformation (not just technical implementation)

Strong understanding of CI/CD, DevOps, and automation practices

Experience working in regulated or large enterprise environments

Familiarity with AIOps or advanced automation strategies, Increased adoption of SLOs and reliability standards

Reduction in high-severity incidents over time

Improved MTTR and operational efficiency

Increased adoption of standardized observability practices

Reduction in reactive, ticket-driven work across teams

Clear alignment between SRE, PSE, and application teams

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · WWC 2021

2:33 min

Advocating for SRE practices within agency environments

Martin Beránek · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all