TELECOMMUTE Sr. SRE (Storage Platforms)

Nasscomm, Inc.
United States
1 day ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours
Job source

Tech stack

Bash Shell Ubuntu (Operating System) CentOS Continuous Integration Linux Disaster Recovery Domain Name System (DNS) Python (Programming Language) Kernel-Based Virtual Machine NetApp Applications OpenStack Red Hat Enterprise Linux
+24 more
Reliability Engineering Ansible Prometheus Systems Integration Ceph (Software) Private Cloud Environment Scripting SUSE Linux System Availability Grafana Backend Git Pure Storage Kubernetes Infrastructure Automation Frameworks Storage Technologies Information Technology Cinder (Software) Data Management Multiaccess Edge Computing Terraform Legacy Systems Nvme Golang

Job description

We are seeking a highly experienced Sr Site Reliability Engineer - Storage Platforms to design, implement, and support Software Defined Storage (SDS) and Kubernetes platforms in a private cloud environment. This role focuses on scalability, resilience, automation, and performance using Infrastructure-as-Code and GitOps practices. This is a deeply technical role requiring expert-level understanding of Software Defined Storage, Kubernetes, and extensive working knowledge on Linux Operating systems. You will also collaborate with platform and SRE teams to maintain secure, performant, and multitenant-isolated services that serve high-throughput, mission-critical applications. Key Responsibilities

  • Design, implement, and operate large-scale Software Defined Storage architectures across private and public cloud regions within ITIL methodology.
  • Deploy and support enterprise storage platforms (Pure Storage, HPE, NetApp) and SDS solutions (Ceph, Longhorn).
  • Build self-service storage workflows for Kubernetes CSI and OpenStack consumers (VM and Baremetal).
  • Develop Infrastructure-as-Code using Ansible, Terraform, Helm and Git, with Python/Bash automation.
  • Implement CI/CD pipelines for infrastructure updates, patching, upgrades, testing, and rollback.
  • Build observability, alerting, and auto-remediation using GitOps and tools such as Prometheus, Loki, and Grafana.
  • Architect and maintain high availability, disaster recovery, and scale-out infrastructure.
  • Develop and review high-level and low-level design documents for storage infrastructure
  • Perform deep troubleshooting across storage, Kubernetes, hypervisors, networking, and Linux systems.
  • Participate in on-call rotations, incident response, and root cause analysis.
  • Collaborate globally on change management, documentation, and operational best practices.

Requirements

  • 6+ years of experience managing enterprise storage and Kubernetes platforms on Linux.
  • Strong hands-on experience with SDS solutions (Ceph, Longhorn) and storage migrations from legacy systems.
  • Experience with block, file, and object storage, including Fibre Channel and IP-based protocols.
  • Experience with NVMe-oF or iSCSI fabrics.
  • Expert knowledge of Kubernetes and Linux systems (Ubuntu, RHEL/CentOS).
  • Proficiency with Infrastructure-as-Code (IaC) (Ansible, Terraform).
  • Strong scripting skills in Python and Bash (Golang (GO) a plus).
  • Strong working knowledge of Enterprise DNS and integrations with Kubernetes
  • Experience operating 24x7 mission-critical production environments.
  • Hands-on experience with KVM hypervisors (Suse Harvester, OpenStack).
  • Strong written and verbal communication skills.
  • Proficiency with Git, CI/CD pipelines, and automated testing frameworks
  • Ability to write technical documentation and contribute to community wikis or knowledge bases.
  • Bachelor’s degree in computer science or equivalent professional experience.

Nice to Have

  • OpenStack Cinder multi-backend administration.
  • Backup platforms (Rubrik).
  • Understanding of CIS/NIST security and infrastructure lifecycle management.
  • ITIL Foundation/advanced certifications in support of ITSM standard methodology.
  • Background in telco, edge cloud, or large enterprise environments.
  • CNCF Certified Kubernetes Administrator (CKA), Certified Kubernetes Security
  • Specialist (CKS) or Red Hat specialist in Ceph Storage Administrator (EX125) certifications.
  • Master’s degree in computer science, IT, Engineering, or a related field preferred;
  • equivalent experience and relevant industry certifications will also be considered.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

1:44 min

Career transition into cloud native and data management

Michael Cade · LIVE

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all