Site Reliability Engineer II (NGPOS Operations Support)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+11 more
Job description
We are seeking a hands-on Site Reliability Engineer II to support a next-generation Point of Sale (NGPOS) platform in a highly visible production environment. Unlike traditional SRE roles focused primarily on automation or platform engineering, this position emphasizes production reliability, incident leadership, operational excellence, and engineering support.
The ideal candidate will lead major incident response efforts, drive Root Cause Analysis (RCA), improve system observability, and collaborate closely with Software Engineering, Platform Engineering, Infrastructure, and Business Operations teams to enhance overall platform reliability.
This role is ideal for someone who enjoys solving complex production issues under pressure while contributing to long-term engineering improvements., * Lead major incident response during production outages
- Serve as Incident Commander during P1/P2 incidents
- Coordinate technical bridge calls
- Communicate outage status to engineering teams and business leadership
- Lead Root Cause Analysis (RCA) activities
- Track corrective actions through completion
- Improve production reliability and system stability
- Enhance monitoring and observability
- Reduce alert fatigue
- Partner with Software Engineering and Platform Engineering teams
- Support retail store deployments
- Develop operational documentation, runbooks, and playbooks
- Participate in after-hours support rotations and maintenance windows
- Improve service health using SLIs and SLOs
Technical Environment
Monitoring & Observability
- Dynatrace
- Azure Monitor
- Log Analytics
- Metrics
- Dashboards
Cloud
- Microsoft Azure
- Google Cloud Platform (GCP)
Containers
- Kubernetes
- Docker
Operating Systems
- Linux
Scripting Languages
- Bash
- Python
Agile Tools
- Jira
Enterprise Environment
- Retail systems
- Point of Sale (POS)
- Production Support
- Hybrid Infrastructure, + Cloud environments
- On-premise infrastructure
- Retail/POS systems
- Observability and Monitoring
- Dynatrace
- Azure Monitor
- Log analysis
- Dashboards
- Metrics
- Linux administration
- Bash and/or Python scripting
- Kubernetes
- Docker
- Microsoft Azure and/or Google Cloud Platform (GCP)
- Agile methodology
- Jira
- Strong communication and collaboration skills
- Ability to work onsite five days per week
- Willingness to travel and participate in on-call rotations
Requirements
Work Authorization: Permanent Residents only. Must be able to convert to full-time without sponsorship. Experience Required: 3+ Years
Nice to Have
- Enterprise Point of Sale (POS) systems
- Retail technology experience
- Automation scripting
- Monitoring optimization
- Runbook creation
- Store technology deployments, * Strong leadership during production incidents
- Excellent troubleshooting and analytical skills
- Effective communication under pressure
- Ownership and accountability
- Experience coordinating multiple engineering teams
- Strong operational discipline
- Continuous improvement mindset
- Passion for reliability engineering
- Excellent documentation skills, Must Have
- Major Incident Management experience
- Experience serving as Incident Commander
- Leading production bridge calls
- Coordinating cross-functional technical teams
- Executive communication during P1/P2 outages
- Root Cause Analysis (RCA)
- Five Whys
- Fishbone Analysis
- Timeline reconstruction
- Corrective action tracking
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.indeed.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Dev Digest 120 - Apple and peers
The Best Job Search Websites of 2025
Find a Developer Job: 12 Best Job Sites For Developers
Highest Paying Tech Companies for Developers