Production Support Engineer (Tier III Operations)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+5 more
Job description
The Tier III Operations Engineer is responsible for maintaining the availability, performance, and reliability of mission-critical systems. This role serves as the highest level of operational support, providing advanced technical troubleshooting and incident resolution for complex infrastructure and application issues. The engineer collaborates with Operations, Engineering, and external vendors to ensure system stability and continuous service availability., * Lead diagnosis and resolution of complex incidents that cannot be resolved by Tier I or Tier II support teams
- Perform advanced troubleshooting across servers, operating systems, networking, databases, storage, virtualization, and enterprise applications
- Participate in incident response during critical outages and coordinate technical recovery efforts
- Monitor system health, performance, and capacity to proactively identify potential issues
- Develop and maintain operational documentation, troubleshooting guides, and standard operating procedures
- Support planned maintenance, system upgrades, patching, and change-management activities
- Collaborate with engineering teams to improve system resiliency, automation, and operational efficiency
- Identify recurring issues and recommend long-term improvements to increase reliability and reduce operational risk
Requirements
- Security+ or CISSP certification
- Strong communication skills with the ability to explain complex technical topics
- Strong understanding of Linux operating systems
- Availability to work extended hours and provide after-hours and on-call support
- Experience supporting and troubleshooting APIs and security appliances such as API gateways
- Ability to analyze system logs, performance metrics, and diagnostic data to quickly isolate technical issues
- Ability to work independently and effectively under pressure during critical production incidents
Preferred Qualifications
- Bachelor’s degree in a STEM field and 5+ years of related experience
- Experience supporting high-availability or mission-critical environments
- Experience with Grafana or similar monitoring tools
- Familiarity with incident management and change-management processes
- Broad knowledge of servers, networking, storage, and virtualization
- Experience with cloud monitoring and logging tools such as Grafana, Prometheus, Promtail, and Loki
- Familiarity with DevOps collaboration tools such as Jira and Confluence
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
Best Paying Jobs in Technology
DevOps Engineer Salary [2023]
Is Software Engineering Over-Saturated?