IT Triage Engineer

First Citizens
United States
20 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Microsoft Windows Applications Architecture Application Performance Management Bash Shell Cyber Security Decision Support Systems Monitoring of Systems Hyper-V Systems Analysis Python (Programming Language) Linux System Administration OpenShift
+13 more
Windows PowerShell Runbook Server Administration Software Engineering Systems Architecture User Environment Management Virtualization Technology Scripting Mttr Technical Debt Kubernetes ArcSight Event Correlation Vmware

Job description

The Senior IT Triage Engineer serves as a critical technical leader responsible for the rapid diagnosis, restoration, and resolution of complex technology incidents across enterprise infrastructure, cloud platforms, applications, and end-user services. This role combines deep technical troubleshooting expertise with a proactive focus on eliminating recurring issues through structured problem management, root cause analysis, and continuous service improvement.

The ideal candidate excels in high-pressure operational environments, drives technical resolution efforts across multiple teams, and leverages automation, observability, and reliability practices to improve service stability and reduce operational risk.

Responsibilities

Key Responsibilities

  • Incident Management, Service Restoration, and System Analysis.
  • Lead technical triage efforts for high-priority incidents and service disruptions.
  • Coordinate cross-functional teams to restore critical business services as quickly as possible.
  • Analyze alerts, logs, monitoring data, and telemetry to identify the source of issues.
  • Serve as a senior escalation point for complex infrastructure, cloud, network, application, and platform incidents.
  • Drive incident bridges, facilitate technical discussions, and maintain clear communication with stakeholders.
  • Problem Management & Root Cause Elimination
  • Lead root cause investigations for recurring or significant incidents.
  • Develop and drive corrective and preventive action plans across technology teams.
  • Track and manage problem records through resolution.
  • Identify systemic issues and technical debt that impact service stability.
  • Attend post-incident reviews and ensure lessons learned are translated into operational improvements.

Reliability & Operational Excellence

  • Continuously improve service availability, resiliency, and operational performance.
  • Partner with engineering teams to improve monitoring, alerting, telemetry, and observability capabilities.
  • Reduce alert noise through event correlation, automation, and process optimization.
  • Drive efforts to improve mean time to detect (MTTD) and mean time to restore service (MTTR).

Automation & Process Improvement

  • Identify opportunities to automate operational workflows and repetitive support activities.
  • Develop runbooks, playbooks, and operational procedures.
  • Collaborate with engineering teams to implement self-healing and automated recovery mechanisms.

Technical Leadership

  • Provide mentorship and guidance to engineers and operational support teams.
  • Influence technical decision-making related to operational readiness and supportability.
  • Act as a trusted advisor for service reliability, supportability, and operational risk management.
  • Participate in change reviews to ensure production readiness and minimize operational impact.

Requirements

Bachelor’s Degree and 8 years of experience in Technical work in Application Development, Server Administration, Information Security, or Engineering OR High School Diploma or GED and 12 years of experience in Technical work in Application Development, Server Administration, Information Security, or Engineering

  • 7+ years of experience in enterprise IT operations, infrastructure engineering, platform operations, or technical support environments.
  • Proven experience managing and resolving critical production incidents across multiple system architectures and infrastructures.
  • Strong background in problem management and root cause analysis methodologies.
  • Experience supporting large-scale enterprise environments.
  • Strong understanding of ITIL service management practices.

Technical Skills

  • Modern application architecture patterns and operations
  • Infrastructure & Platforms
  • Windows and Linux administration
  • Virtualization platforms (VMware, Hyper-V, etc.)
  • Storage and backup technologies
  • Containers and orchestration platforms (OpenShift)

Monitoring & Observability

  • Enterprise monitoring platforms
  • Log aggregation and analysis tools
  • Application performance monitoring (APM)
  • Operational dashboards and telemetry systems
  • Event management solutions

Automation

  • PowerShell, Python, Bash, or equivalent scripting
  • Workflow automation platforms
  • API integration and orchestration
  • Runbook development

Key Competencies

  • Exceptional troubleshooting and analytical skills.
  • Strong sense of ownership and accountability.
  • Ability to perform effectively during major incidents and high-pressure situations.
  • Excellent collaboration and stakeholder management skills.
  • Strong written and verbal communication.
  • Data-driven decision making.
  • Systems thinking and problem-solving mindset.
  • Continuous improvement orientation.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:40 min

Managing containerized infrastructure with Podman Desktop

Cedric Clyburn Cedric Clyburn +1 · WWC 2025

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · WWC 2021

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

1:41 min

Parallels between cloud and legacy infrastructure lock-ins

Björn Stahl Björn Stahl · WWC 2024

3:07 min

Establishing service level agreements directly for internal platforms

Pawel Piwosz · LIVE

Videos

See all

Related articles

See all