Site Reliability Engineer II (NGPOS Operations Support)

Hudson
Cincinnati, OH, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$104,000.0 - $108,160.0
Working hours
Shift work
Job source

Tech stack

Agile Methodology JIRA Microsoft Azure Bash Shell Cloud Computing Monitoring of Systems Python (Programming Language) Linux System Administration Log Analysis Reliability Engineering Retail Software Runbook
+11 more
Shell Script Software Engineering Scripting Google Cloud Cloud Platform System Kubernetes Atlassian Tools Performance Monitor Hardware Infrastructure Dynatrace Docker

Job description

We are seeking a hands-on Site Reliability Engineer II to support a next-generation Point of Sale (NGPOS) platform in a highly visible production environment. Unlike traditional SRE roles focused primarily on automation or platform engineering, this position emphasizes production reliability, incident leadership, operational excellence, and engineering support.

The ideal candidate will lead major incident response efforts, drive Root Cause Analysis (RCA), improve system observability, and collaborate closely with Software Engineering, Platform Engineering, Infrastructure, and Business Operations teams to enhance overall platform reliability.

This role is ideal for someone who enjoys solving complex production issues under pressure while contributing to long-term engineering improvements., * Lead major incident response during production outages

  • Serve as Incident Commander during P1/P2 incidents
  • Coordinate technical bridge calls
  • Communicate outage status to engineering teams and business leadership
  • Lead Root Cause Analysis (RCA) activities
  • Track corrective actions through completion
  • Improve production reliability and system stability
  • Enhance monitoring and observability
  • Reduce alert fatigue
  • Partner with Software Engineering and Platform Engineering teams
  • Support retail store deployments
  • Develop operational documentation, runbooks, and playbooks
  • Participate in after-hours support rotations and maintenance windows
  • Improve service health using SLIs and SLOs

Technical Environment

Monitoring & Observability

  • Dynatrace
  • Azure Monitor
  • Log Analytics
  • Metrics
  • Dashboards

Cloud

  • Microsoft Azure
  • Google Cloud Platform (GCP)

Containers

  • Kubernetes
  • Docker

Operating Systems

  • Linux

Scripting Languages

  • Bash
  • Python

Agile Tools

  • Jira

Enterprise Environment

  • Retail systems
  • Point of Sale (POS)
  • Production Support
  • Hybrid Infrastructure, + Cloud environments
  • On-premise infrastructure
  • Retail/POS systems
  • Observability and Monitoring
  • Dynatrace
  • Azure Monitor
  • Log analysis
  • Dashboards
  • Metrics
  • Linux administration
  • Bash and/or Python scripting
  • Kubernetes
  • Docker
  • Microsoft Azure and/or Google Cloud Platform (GCP)
  • Agile methodology
  • Jira
  • Strong communication and collaboration skills
  • Ability to work onsite five days per week
  • Willingness to travel and participate in on-call rotations

Requirements

Work Authorization: Permanent Residents only. Must be able to convert to full-time without sponsorship. Experience Required: 3+ Years

Nice to Have

  • Enterprise Point of Sale (POS) systems
  • Retail technology experience
  • Automation scripting
  • Monitoring optimization
  • Runbook creation
  • Store technology deployments, * Strong leadership during production incidents
  • Excellent troubleshooting and analytical skills
  • Effective communication under pressure
  • Ownership and accountability
  • Experience coordinating multiple engineering teams
  • Strong operational discipline
  • Continuous improvement mindset
  • Passion for reliability engineering
  • Excellent documentation skills, Must Have
  • Major Incident Management experience
  • Experience serving as Incident Commander
  • Leading production bridge calls
  • Coordinating cross-functional technical teams
  • Executive communication during P1/P2 outages
  • Root Cause Analysis (RCA)
  • Five Whys
  • Fishbone Analysis
  • Timeline reconstruction
  • Corrective action tracking

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

2:50 min

Introduction and the value of runbooks

Hila Fish · WWC 2023

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:32 min

Structuring automated incident workflows between runbooks and raw models

Aram Hakobyan Aram Hakobyan +1 · WWC Europe 2026

Videos

See all

Related articles

See all