> Markdown version of [/jobs/ext/1727059-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1727059-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** CRC Insurance Services Inc - **Location:** Charlotte, NC, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Microsoft Azure, Configuration Management Databases, Data Deduplication, Data Synchronization, Noise Reduction, Github, Monitoring of Systems, Infrastructure as a Service (IaaS), Platform as a Service (PAAS), Reliability Engineering, Site Reliability Engineering Practices, Runbook, Wireless Paging Systems, ServiceNow IT Service Management, Mttr, Performance Monitor, Bicep, Terraform, Dynatrace, Pagerduty, Servicenow - **Published:** July 16, 2026 - **Apply:** https://dejobs.org/x/x/8C6A3D83CDBB4B978F3125A6837D8449/job/ ## About the Role * 7+ years in Site Reliability Engineering, monitoring, or production engineering * Proven experience in a technical leadership or lead engineer role * Deep hands-on experience with: * Dynatrace (or equivalent observability platforms) * Microsoft Azure (IaaS, PaaS, networking, identity) * ServiceNow ITSM / ITOM (incident, event management, CMDB) * Demonstrated ability to: * Design and lead enterprise monitoring/SRE architectures * Drive platform and tooling decisions * Integrate observability, ITSM, and paging solutions, * Experience leading SRE or observability transformation initiatives * Strong expertise with Dynatrace-ServiceNow integrations * Experience modernizing or consolidating paging/on-call tooling * Familiarity with: * Azure-based SRE tooling or AI-assisted operations * Automation frameworks (GitHub Actions, Runbooks, etc.) * Infrastructure as Code (Terraform, ARM, Bicep) Success Metrics * Reduction in alert noise and unnecessary paging * Improved incident routing accuracy and MTTR * Increased adoption of self-healing and automated workflows * Strong alignment between monitoring, CMDB, and service ownership * Enterprise-wide adoption of SRE and monitoring standards, We seek passionate individuals who thrive in a fast-paced, collaborative environment. If you value integrity and are driven to succeed, CRC Group is the place for you. ## Description We are seeking a Lead Site Reliability & Environment Monitoring Engineer to establish and evolve our enterprise observability and monitoring strategy across cloud and application platforms. This is a full-time leadership role responsible for owning monitoring design, driving platform decisions, and guiding engineering teams toward modern SRE practices. This individual will act as the technical authority for monitoring and alerting, shaping how signals from Dynatrace flow into ServiceNow and enterprise messaging/paging platforms, and enabling a shift toward automated, intelligent, and self-healing operations., Strategic Leadership & Decision-Making * Define and own the enterprise monitoring and SRE observability strategy * Serve as the subject matter expert for Dynatrace, ServiceNow integration, and alerting architecture * Evaluate and recommend tooling, integration patterns, and platform direction * Drive decisions on alerting philosophy, noise reduction, and signal quality improvement Platform Ownership & Architecture * Architect and standardize end-to-end monitoring and SRE pipelines: * Dynatrace * ServiceNow incident lifecycle * Alert correlation, deduplication, and prioritization * Integration with paging systems (PagerDuty, SMS, voice, Teams) * Establish best practices for: * Event ingestion and enrichment * Incident routing and automated assignment * Integration with CMDB and service mapping Site Reliability Engineering (SRE) Leadership * Lead adoption of SRE principles, including: * SLIs, SLOs, and error budgets * Reliability engineering practices across services * Proactive monitoring and resilience design * Champion a shift from reactive operations to proactive reliability engineering * Influence application and platform teams to build observable, resilient systems by design Automation & Self-Healing Enablement * Drive development of automated remediation and self-healing capabilities * Leverage Dynatrace workflows, Azure services, and automation frameworks to: * Reduce manual incident handling * Eliminate repeatable operational tasks * Minimize unnecessary paging ServiceNow & Observability Integration Leadership * Own integration between Dynatrace and ServiceNow ITSM/ITOM, including: * Incident, Event Management, and CMDB alignment * Service mapping and dependency visibility * Governance for application/service tagging * Define standards for: * Automated incident creation and resolution * Priority assignment and routing logic * Monitoring-to-ITSM data synchronization Team Leadership & Cross-Functional Influence * Provide technical leadership and mentorship across SRE, platform, and application teams * Act as a central point of coordination between engineering, cloud, and ITSM teams * Lead workshops and working sessions to: * Drive monitoring standardization * Align teams on reliability practices * Influence upstream architectural decisions Operational Excellence * Establish KPIs and drive improvement in: * Incident response and resolution times * Alert quality and paging effectiveness * Monitoring coverage across critical services * Provide leadership with clear visibility into service health and reliability trends ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Back(end) to the Future: Embracing the continuous Evolution of Infrastructure and Code](https://www.wearedevelopers.com/videos/440-back-end-to-the-future-embracing-the-continuous-evolution-of-infrastructure-and-code) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [Events like RSAC Get You CISOs. Developers Decide What Actually Gets Deployed.](https://www.wearedevelopers.com/magazine/693-events-like-rsac-get-you-cisos-developers-decide-what-actually-gets-deployed) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated)