Application Support Level 3 (L3 Support / Site Reliability Engineer)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+18 more
Job description
We are seeking an Application Support Level 3 (L3 Support / Site Reliability Engineer) to support a next-generation Point of Sale (POS) platform. This role is embedded within the engineering team and works closely with developers throughout the Software Development Life Cycle (SDLC). The ideal candidate is a strong problem solver with excellent troubleshooting skills, experience handling production incidents, and a passion for improving system reliability. Candidates should be comfortable working in a fast-paced production environment, participating in on-call support, traveling for new store deployments, and collaborating with cross-functional teams to ensure high system availability. Required Qualifications, * Provide Level 3 production support for enterprise Point of Sale applications.
- Monitor production systems and respond quickly to incidents and service interruptions.
- Lead and coordinate Major Incident Management activities for critical production issues.
- Drive incidents through resolution while meeting established SLAs.
- Perform Root Cause Analysis (RCA) and implement corrective actions to prevent recurring issues.
- Troubleshoot application, infrastructure, cloud, and store-level production issues.
- Partner with developers, infrastructure teams, and business stakeholders to resolve complex problems.
- Test application and infrastructure changes before production deployment.
- Participate in application design, development, coding, and unit testing activities following SDLC best practices.
- Develop functional specifications and implementation plans for application enhancements.
- Create estimates and work plans for development and deployment activities.
- Execute break/fix activities, service requests, and production support tasks.
- Maintain operational procedures, scripts, and documentation.
- Continuously improve software delivery processes and engineering standards.
- Participate in Agile ceremonies and support sprint planning, refinement, and testing activities.
- Build strong partnerships across engineering, infrastructure, and business teams.
- Lead projects and mentor junior team members when needed.
Requirements
- Bachelor’s degree in Computer Science, Information Systems, or a related field (or equivalent experience).
- Experience in Application Support, Site Reliability Engineering (SRE), Operations, or Software Engineering.
- Strong production support and troubleshooting experience.
- Experience managing Major Incidents (P1/P2) and coordinating incident response.
- Knowledge of Root Cause Analysis (RCA) methodologies such as 5 Whys, Fishbone Analysis, and Timeline Reconstruction.
- Experience monitoring applications and infrastructure using enterprise monitoring tools.
- Understanding of Java applications; Go experience is a plus.
- Exposure to SQL or NoSQL databases such as MongoDB.
- Experience with scripting using Shell, PowerShell, or Python is preferred.
- Familiarity with ServiceNow, Jira, Confluence, Splunk, Grafana, Dynatrace, or similar monitoring and ITSM tools.
- Exposure to Point of Sale (POS) systems is preferred.
- Understanding of Agile methodologies; Scrum Master experience is a plus.
- Knowledge of PCI-DSS or payment card security standards is a plus.
- Excellent verbal and written communication skills.
- Ability to work onsite five days per week.
- Willingness to travel for store deployments, pilots, and production rollouts.
- Ability to participate in a rotating 24x7 on-call schedule., * Experience supporting enterprise Point of Sale (POS) environments.
- Knowledge of Java and Go programming.
- Experience with SQL and MongoDB.
- Experience using monitoring and observability tools such as Splunk, Dynatrace, or Grafana.
- Familiarity with cloud platforms and hybrid environments.
- Experience with Agile and Scrum practices.
- Understanding of payment processing and PCI compliance.
Top Skills Required
- Major Incident Management with hands-on experience leading P1/P2 production incidents.
- Strong Root Cause Analysis (RCA) and Problem Management experience.
- Production troubleshooting across application, cloud, infrastructure, and POS environments.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
What’s the Difference between a Junior, Mid, and Senior Developer?
Is Software Engineering Over-Saturated?
Software Developer Salary in Switzerland [2023]