Digital Support Analyst

Central Business Solutions Inc
Frisco, TX, United States
26 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Query Performance Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Confluence JIRA Microsoft Azure Batch Processing Cloud Computing Databases DevOps Digital Technology
+15 more
Release Management Runbook Salesforce.Com Datadog Google Cloud Cloud Platform System Grafana Backend Low Latency Atlassian Tools Performance Monitor Bitbucket Splunk Dynatrace Servicenow

Job description

We are seeking a highly motivated Digital Support Analyst (Level 1.5/Level 2) to help establish a modern, structured digital support and services model. This role will serve as a critical triage and diagnostic layer between the service desk (L1) and engineering teams (L3), transforming support from a reactive intake function into a proactive, insight-driven operations capability.

Reporting to the Director, Digital Service Management & Operations, this role will focus on incident triage, monitoring, troubleshooting, and release-aware diagnostics across a complex, integrated application and cloud environment.


Context & Mission

The current support model operates as a broad, centralized intake layer covering end-user support, applications, and cloud infrastructure, but lacks:

ยท Dedicated monitoring (โ€œeyes-on-glassโ€)

ยท Structured triage and diagnostic capability

ยท Deep technical ownership at the L1/L2 layers

This role is designed to:

ยท Introduce structured L1.5/L2 triage and investigation

ยท Improve incident quality and escalation effectiveness

ยท Enable proactive monitoring and operational stability

ยท Support a high-frequency release environment


Key Responsibilities

  1. Incident Triage & Technical Investigation (Core Focus)

ยท Act as the primary L1.5/L2 triage layer, analyzing alerts, incidents, and system anomalies.

ยท Perform initial troubleshooting and diagnostics before escalation to L3 teams.

ยท Develop structured problem statements with actionable insights for engineering teams.

ยท Interpret and analyze:

o API responses and endpoint behavior

o Application errors (e.g., HTTP 500s)

o Latency and performance degradation

o Infrastructure and database indicators (CPU, memory, query performance)

ยท Participate in incident bridge calls, providing real-time technical insights.


  1. Monitoring & โ€œEyes-on-Glassโ€ Operations

ยท Support the development of a dedicated monitoring capability.

ยท Continuously monitor systems using observability/APM tools (e.g., Datadog, Google Cloud Platform, Idera).

ยท Correlate alerts and identify patterns, anomalies, and early warning signals.

ยท Shift operations from reactive escalation * proactive issue detection and prevention.


  1. Cross-Functional Troubleshooting

ยท Investigate issues across integrated front-end, back-end, and cloud platforms (e.g., Google Cloud Platform environments).

ยท Perform cross-stack analysis, recognizing that issues may span multiple systems.

ยท Provide context-aware insights (e.g., dependency-related failures like external platform issues).


  1. Release-Aware Support & QA Mindset

ยท Collaborate with the Release Manager and DevOps teams on frequent (multi-weekly) releases.

ยท Correlate incidents with:

o Recent deployments

o Configuration changes

o Feature rollouts

ยท Apply a QA-oriented mindset to validate system behavior post-release.

ยท Support hypercare and post-release monitoring activities.


  1. Runbooks, Knowledge & Process Maturity

ยท Execute and continuously improve runbooks and standard operating procedures (SOPs).

ยท Document:

o Troubleshooting steps

o Known issues

o Root cause indicators

ยท Drive standardization of triage and escalation processes.

ยท Contribute to problem management lifecycle and post-incident reviews (PIRs).


  1. Service Management & ITIL Alignment

ยท Operate within ServiceNow (incident, change, and problem management) and Atlassian (JIRA, Confluence, Bitbucket).

ยท Improve:

o Ticket quality

o Categorization and prioritization

o Escalation readiness

ยท Support CAB readiness activities by providing operational insights.

ยท Contribute to defining an L1-L2-L3 escalation framework.


  1. Automation, AI & Continuous Improvement

ยท Leverage emerging AIOps and automation capabilities for:

o Alert correlation

o Root cause identification

ยท Identify opportunities to reduce manual triage effort.

ยท Collaborate on improving integration between:

o ITSM (ServiceNow)

o Monitoring/observability tools

Requirements

ยท 3-6 years of experience in application support, production support, or digital operations.

ยท Strong troubleshooting experience in complex, integrated environments.

ยท Hands-on experience with:

o Monitoring/APM tools (Datadog, Dynatrace, Splunk, etc.)

o Incident management platforms (ServiceNow preferred)

ยท Ability to interpret:

o Logs, APIs, and system metrics

o Performance and error patterns

ยท Solid understanding of ITIL processes (Incident, Problem, Change).


Preferred Qualifications

ยท Experience in cloud environments (Google Cloud Platform, Azure, or AWS).

ยท Exposure to CI/CD pipelines and DevOps practices.

ยท Familiarity with:

o CRM platforms (e.g., Salesforce, Dynamics)

o Job orchestration/batch processing tools

ยท Experience in high-availability, customer-facing digital systems.


Key Competencies

ยท Strong analytical and diagnostic mindset

ยท Ability to move beyond escalation to true problem investigation

ยท Curiosity and ownership in understanding root causes

ยท Effective communication in high-pressure incident scenarios

ยท Ability to work across applications, infrastructure, and QA domains

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role โ€” technically off-topic, practically not.

1:06 min

Developer experience and project variety at scale

Alexandra Petri ยท WWC 2023

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum ยท WWC Europe 2026

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 ยท LIVE

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein ยท LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark ยท LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 ยท LIVE

Videos

See all

Related articles

See all