NGPOS Ops Onsite Support Engineer

BTC Inc.
Blue Ash, OH, United States
2 days ago
Apply on www.thejobnetwork.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

JIRA Microsoft Azure Bash Shell Cloud Computing Continuous Integration Linux DevOps Python (Programming Language) Reliability Engineering Scripting Kubernetes Dynatrace
+1 more
Docker

Job description

Specialize in ensuring the reliability, resilience, and recoverability of enterprise services and systems across cloud, on-prem, and store environments. Serve as the SRE voice within incident and problem management, leading major incident response, driving root cause analysis, and partnering with Software Engineers and platform teams to reduce recurrence and improve system health in production., Role Functions: Act as the SRE lead/participant in Major Incident Management: joining bridges, driving triage, coordinating cross-team response, and communicating status to stakeholders during active incidents Own or contribute to Problem Management: conducting RCAs, identifying systemic issues, and tracking corrective actions to completion in partnership with engineering and the platform SRE team Collaborate with business partners and engineering teams to define and monitor key SLOs/SLIs, and use them to prioritize reliability work Improve traceability, observability, and retrievability of system behavior in production using tools such as Dynatrace and Azure Build/tune monitoring and alerting to reduce noise, catch issues earlier, and speed up diagnosis across cloud, on-prem, and store environments Partner with the dedicated platform SRE team to escalate systemic/longer-term reliability gaps uncovered during incident and problem work Author clear playbooks, runbooks, and postmortem documentation for use by the broader engineering organization Participate in an off-hours on-call rotation and periodic off-hours work during major incidents or maintenance windows Demonstrate the company's core values of respect, honesty, integrity, diversity, inclusion, and safety

Note to Vendors Top 3 skills Major Incident Management experience has actually run or driven a bridge/war-room for a P1/P2 outage, not just 'participated.' This is the core of the role; everything else is supporting it. Ask for a specific incident they led end-to-end. RCA/Problem Management rigor can articulate a structured RCA methodology (5-whys, fishbone, timeline reconstruction) and, more importantly, examples where they drove a corrective action to closure (not just wrote a doc that nobody actioned). Production troubleshooting under pressure across a mixed stack comfortable reading logs/dashboards (Dynatrace-type tooling) and reasoning about cloud + on-prem + store/POS systems simultaneously, since a real incident here will span all three. Are there any automatic disqualifiers No on-call/off-hours availability Cant commit to 5 days/week in Blue Ash, OH No travel availability for go-lives/pilots Pure DevOps/CI-CD background with zero incident-response exposure

Requirements

3+ years of experience in Site Reliability Engineering, incident management, or production support for enterprise systems Experience leading or participating in major incident management (incident commander, bridge/war-room facilitation, executive communication during outages) Experience with problem management and root cause analysis (RCA), including driving corrective/preventive actions to closure Solid understanding of observability and monitoring concepts and tooling (e.g., Dynatrace, Azure Monitor) to detect, triage, and diagnose production issues Working knowledge of Linux and scripting (e.g., BASH, Python) for troubleshooting and diagnostic automation Familiarity with Kubernetes, Docker, and cloud platforms (Azure, GCP) sufficient to troubleshoot and reason about system behavior in production Experience working in an Agile environment, tracking work and metrics via a platform such as Jira Strong communication skills able to translate technical incident details for both engineering teams and business stakeholders under time pressure In office, Blue Ash, OH 5 days a week Willingness to travel and provide onsite support for new go-lives and pilots Familiarity with Point of Sale Systems in Enterprise environments

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.thejobnetwork.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:08 min

Shifting software delivery bottlenecks to operations and incident response

Milin Desai Milin Desai +1 · World Congress 2026 Europe

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:06 min

Developer experience and project variety at scale

Alexandra Petri · World Congress 2023

5:47 min

Integrating user stories and test automation via Jira tools

Christoph Ruggenthaler · LIVE

Videos

See all

Related articles

See all