Site Reliability Engineer

DivIHN Integration
Corning, NY, United States
17 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Agile Methodology Amazon Web Services Application Configuration Access Protocols Systems Engineering Microsoft Azure Bash Shell Common ISDN Application Programming Interface (CAPI) Cloud Computing Continuous Integration Linux DevOps
+20 more
Python (Programming Language) Linux System Administration Octopus Deploy Performance Tuning Scrum Methodology Reliability Engineering Site Reliability Engineering Practices Software Engineering Data Logging Scripting Google Cloud Cloud Platform System High Performance Computing Reliability of Systems Git Kubernetes Infrastructure Automation Frameworks Information Technology Deployment Automation Rancher

Job description

Schedule: Candidates can live anywhere in the US but must be able to work 8 AM - 5 PM/9 AM - 6 PM EST Only W2 candidates are eligible for this position. Third-party or C2C candidates will not be considered., * Client is looking for an experienced contract Site Reliability Engineer who can strengthen our team’s platform engineering and operational capabilities. You will play a key role in supporting Kubernetes infrastructure managed through Rancher, improving system reliability and automation, and advancing infrastructure-as-code and GitOps practices across our environment., * Platform Operations: Maintain and enhance Kubernetes platforms across on-premises and cloud environments, ensuring reliability, scalability, and operational efficiency.

  • Cluster Management: Support provisioning, upgrades, troubleshooting, and lifecycle management of Kubernetes clusters managed through Rancher.
  • Linux Systems Administration: Provide deep technical expertise in Linux-based systems, including performance tuning, troubleshooting, automation, and operational support.
  • Infrastructure as Code: Develop and maintain infrastructure-as-code solutions to standardize and automate platform deployment and management, with a preference for Cluster API (CAPI)-based approaches.
  • GitOps and Deployment Automation: Support and improve GitOps workflows using ArgoCD to manage cluster and application configuration in a consistent, auditable manner.
  • Collaboration: Work closely with developers, scientists, and infrastructure teams to deliver reliable platform services and translate operational needs into sustainable engineering solutions.
  • Continuous Improvement: Identify opportunities to improve platform resilience, observability, security, and maintainability through automation and modern SRE practices.

Requirements

  • Bachelors is preferred, but not required.
  • Minimum of 5 years professional experience in site reliability engineering, platform engineering, DevOps, or systems engineering roles.
  • Candidates must have 5 years strong system admininstration with Linux! Rancher for Kubernetes experience is a must., * Join Client’s Model Operations and Deployment Engineering team at their flagship research facility, where your work will directly support groundbreaking materials science innovations. In this role, you will help maintain, enhance, and evolve the Kubernetes platforms that enable scientific and engineering teams to deploy, operate, and scale critical applications across both on-premises and cloud environments., * BS in Computer Science, Software Engineering, Information Technology, or related field preferred; or equivalent professional experience., * 5+ years of professional experience in site reliability engineering, platform engineering, DevOps, or systems engineering roles.
  • Hands-on experience operating and supporting Kubernetes platforms in production environments.
  • Strong experience managing Kubernetes clusters in both on-premises and cloud-based environments.
  • Strong Linux systems administration skills, including troubleshooting, scripting, networking, and system performance analysis.
  • Experience with Rancher for Kubernetes cluster management and platform operations.
  • Experience implementing infrastructure-as-code solutions for platform provisioning and lifecycle management.
  • Demonstrated success working in Agile teams (Scrum, Kanban).

Technical Skills:

  • Kubernetes: Cluster operations, upgrades, networking, storage, troubleshooting, and workload support.
  • Platform Management: Rancher or similar Kubernetes management platforms.
  • Linux: Advanced administration of Linux/Unix systems.
  • Infrastructure as Code: Strong IaC experience; Cluster API (CAPI) preferred.
  • GitOps/CI-CD: ArgoCD, Git version control, and deployment automation practices.
  • Scripting/Automation: Bash, Python, or similar scripting languages for automation and operational tooling.

Preferred:

  • Experience with hybrid infrastructure spanning on-premises and public cloud platforms (AWS, Azure, Google Cloud Platform).
  • Experience with Kubernetes ecosystem tooling for observability, logging, monitoring, and alerting.
  • Familiarity with security best practices for Kubernetes and Linux platforms.
  • Experience supporting scientific research environments, high-performance computing, or computational science workflows.
  • Knowledge of CI/CD pipeline development and platform automation patterns.

About the company

DivIHN, the ‘‘IT Asset Performance Services’’ organization, provides Professional Consulting, Custom Projects, and Professional Resource Augmentation services to clients in the Mid-West and beyond. The strategic characteristics of the organization are Standardization, Specialization, and Collaboration.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

Videos

See all

Related articles

See all