Site Reliability Engineer (SRE) - Only

Saransh Inc
Houston, TX, United States
about 1 month ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Working hours
Regular working hours

Tech stack

Link Aggregation (Ethernet) Automation of Tests Backup Devices Configuration Management Desktop Virtualization Disaster Recovery Failover Clustering Hyper-V System Center Operations Management Windows Servers Performance Tuning Windows PowerShell
+23 more
Release Management Reliability Engineering Site Reliability Engineering Practices Prometheus Runbook System Center Virtual Machine Management Virtual Local Area Networks Virtual Machine Manager Private Cloud Environment Data Logging Cloud Monitoring System Availability Grafana Mttr HybridCloud Infrastructure as Code (IaC) Infrastructure Automation Frameworks Information Technology Performance Monitor Patch Management Veeam Splunk Nvme

Job description

We are seeking an experienced Site Reliability Engineer (SRE) Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise with modern Site Reliability Engineering practices to improve platform reliability, scalability, performance, automation, and operational excellence. The ideal candidate will be responsible for ensuring platform reliability through infrastructure automation, OS upgrades, proactive monitoring, incident response, capacity planning, and continuous service improvement while collaborating with infrastructure, security, and application teams., * Operate enterprise-scale private cloud infrastructure built on Microsoft Hyper-V.

  • Optimize, and support highly available VDI environments on Hyper-V.
  • Improve platform reliability, availability, scalability, and resiliency by applying SRE principles and engineering best practices.
  • Disaster recovery, backup, patch management, and business continuity strategies.
  • Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational metrics for critical infrastructure services.
  • Automate infrastructure provisioning, configuration management, and operational workflows using PowerShell and Infrastructure as Code (IaC) principles wherever applicable.
  • Manage Hyper-V Failover Clusters, host lifecycle, storage, networking, and capacity to ensure high availability and business continuity.
  • Develop proactive monitoring, alerting, logging, and observability capabilities to detect and prevent service degradation.
  • Lead incident response for infrastructure-related outages, perform root cause analysis (RCA), and implement preventive actions through post-incident reviews.
  • Perform capacity planning, performance tuning, and resource optimization across Hyper-V clusters and VDI platforms.
  • Support infrastructure migration initiatives including P2V, V2V, workload modernization, and private cloud transformations.
  • Collaborate closely with Security, Networking, Platform Engineering, and Application teams to improve platform reliability and operational efficiency.
  • Develop and maintain technical documentation, architecture diagrams, operational runbooks, automation scripts, and standard operating procedures.
  • Mentor junior engineers and promote SRE culture, automation, and operational best practices across the team., Job Category: Technical Job Description: As a Site Reliability Engineer, you will be responsible for: Operational Excellence & Incident Management - Maintain and monitor prod…
  • 1 month ago

Requirements

Site Reliability Engineering

  • Strong understanding of Site Reliability Engineering principles and operational excellence.
  • Experience with infrastructure reliability, service availability, resiliency, and performance optimization.
  • Storage Space Direct and failover clustering technical expertise. (Storage Spaces Direct enables you to build highly available, software-defined storage by pooling local disks (SSDs, NVMe drives, and HDDs) across multiple Windows Server nodes in a cluster. Instead of relying on an external SAN, S2D uses the servers’ local storage to create a resilient shared storage pool)
  • Experience managing production-critical infrastructure environments with high availability requirements.
  • Experience with incident management, problem management, RCA, and continuous operational improvement.
  • Knowledge of monitoring, observability, alerting, and performance management.

Microsoft Hyper-V (Core Expertise)

  • Deep hands-on expertise in Microsoft Hyper-V architecture, deployment, administration, troubleshooting, and optimization.
  • Extensive experience in operating enterprise private cloud environments on Hyper-V.
  • Strong experience supporting enterprise-scale VDI deployments on Hyper-V.
  • Hyper-V Failover Clustering and high-availability architecture.
  • Storage integration including SAN, NAS, Storage Spaces Direct (S2D), Cluster Shared Volumes (CSV), and storage optimization.
  • Networking within Hyper-V environments including virtual switches, VLANs, NIC Teaming, QoS, and network performance tuning.
  • System Center Virtual Machine Manager (SCVMM).

Automation & Platform Engineering

  • Strong PowerShell scripting and automation experience.
  • Experience automating infrastructure deployment, operational tasks, health checks, and reporting.
  • Familiarity with Infrastructure as Code concepts and configuration management.
  • Experience developing reusable operational tooling to improve reliability and reduce manual effort.

Preferred Skills

  • Windows Server 2016/2019/2022 administration.
  • Experience with backup and disaster recovery solutions such as Veeam, Altaro, or native Hyper-V Replica.
  • Exposure to hybrid cloud and private cloud platforms.
  • Familiarity with monitoring and observability platforms such as SCOM, Azure Monitor, Prometheus, Grafana, Splunk, or similar tools.
  • Experience supporting enterprise VDI environments.
  • Understanding of ITIL Incident, Problem, Change, and Release Management.
  • Experience working in regulated industries such as Banking or Financial Services., * 6+ years of infrastructure engineering experience with at least 4+ years of hands-on Microsoft Hyper-V administration.
  • Demonstrated experience operating mission-critical enterprise infrastructure with high availability and reliability requirements.
  • Proven experience implementing automation to reduce operational overhead and improve service reliability.
  • Experience supporting enterprise private cloud and VDI environments.
  • Experience participating in incident response, root cause analysis, and continuous service improvement initiatives.
  • Microsoft certifications such as Microsoft Certified: Windows Server Hybrid Administrator Associate or equivalent are desirable.
  • Experience in Banking or Financial Services environments is advantageous.

Success Measures

  • Improved platform availability and reliability.
  • Reduced infrastructure incidents through automation and proactive monitoring.
  • Improved Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR).
  • Increased infrastructure automation and operational efficiency.
  • Consistent achievement of service reliability objectives and operational KPIs.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · World Congress 2021

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

3:07 min

Establishing service level agreements directly for internal platforms

Pawel Piwosz · LIVE

4:01 min

Implementing the barbell strategy and focusing on recovery time

Jan de Vries Jan de Vries · World Congress 2026 Europe

9:06 min

Questions on career paths and continuous delivery orchestration platforms

Zan Markan Zan Markan · LIVE

Videos

See all

Related articles

See all