Technical Operations Lead

First Citizens
United States
6 days ago
Apply on diversityjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours

Tech stack

Amazon Web Services Application Performance Management Microsoft Azure Cloud Computing Cloud Engineering Continuous Integration DevOps Monitoring of Systems Log Analysis Reliability Engineering Site Reliability Engineering Practices Software Engineering
+9 more
Google Cloud Mttr Software Troubleshooting Kubernetes Information Technology Performance Monitor Splunk AIOps Dynatrace

Job description

We are seeking a highly skilled and motivated Senior Site Reliability Engineer (SRE) to join our Enterprise Observability and Site Reliability Engineering team. This role is responsible for improving the reliability, availability, performance, and operational excellence of critical enterprise platforms and applications.

The ideal candidate combines strong software engineering and infrastructure expertise with a passion for automation, observability, and operational excellence. You will partner closely with application development, cloud engineering, infrastructure, security, and production support teams to establish reliability standards, implement monitoring strategies, and drive proactive risk reduction across the technology landscape., Reliability Engineering

  • Design, implement, and maintain highly available, resilient, and scalable technology platforms.
  • Manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets across critical services.
  • Lead reliability reviews and identify opportunities to improve system stability and performance.
  • Drive root cause analysis and corrective actions for high-severity incidents during the problem process.

Observability & Monitoring

  • Serve as a subject matter expert for enterprise observability platforms, including Dynatrace and related monitoring technologies.
  • Develop monitoring, alerting, synthetic testing, and dashboard strategies.
  • Improve visibility into application health, infrastructure performance, user experience, and business transaction monitoring.
  • Partner with engineering teams to embed observability practices throughout the software development lifecycle.

Automation & Platform Engineering

  • Identify and eliminate operational toil through automation.
  • Develop scripts, tooling, integrations, and self-service capabilities.
  • Enhance CI/CD processes to improve deployment reliability and operational efficiency.
  • Support infrastructure-as-code and automation-first operating models.

Incident Management & Operational Excellence

  • Participate in major incident response and problem management activities.
  • Establish and improve operational runbooks, standards, and best practices.
  • Drive reduction in Mean Time to Detect (MTTD) and Mean Time to Restore (MTTR).
  • Implement proactive monitoring and predictive alerting capabilities.

Leadership & Collaboration

  • Provide technical leadership and mentoring to engineers across the organization.
  • Influence reliability and observability strategy at the enterprise level.
  • Collaborate with application owners, cloud teams, security teams, and executive stakeholders.
  • Champion a culture of ownership, accountability, and continuous improvement.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, Information Technology, or equivalent experience.
  • 7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Infrastructure Engineering, or related disciplines.
  • Strong experience supporting mission-critical production environments.
  • Hands-on experience with observability and monitoring platforms such as Dynatrace, Splunk, Databahn or similar technologies.
  • Strong understanding of application performance monitoring (APM), distributed tracing, log analytics, and infrastructure monitoring.
  • Experience with cloud platforms such as AWS, Azure, or Google Cloud Platform.
  • Proficiency with scripting and automation.
  • Strong troubleshooting, analytical, and problem-solving skills.
  • Large financial institution experience in a complex environment.

Preferred Qualifications

  • Deep expertise with Dynatrace platform administration and implementation.
  • Experience in enterprise-scale financial services or highly regulated environments.
  • Knowledge of cloud-native observability patterns and OpenTelemetry.
  • Experience implementing SRE practices including SLOs, Error Budgets, reliability reviews, and operational readiness assessments.
  • Familiarity with enterprise event management and AIOps platforms.
  • Certifications in cloud platforms, Dynatrace, Kubernetes, or related technologies.
  • Experience leading large-scale monitoring transformation initiatives.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on diversityjobs.com
Prepare application

Good distractions

Loading talks and stories from around this role…