> Markdown version of [/jobs/ext/2835496-lead-infrastructure-site-reliability-engineering](https://www.wearedevelopers.com/jobs/ext/2835496-lead-infrastructure-site-reliability-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Infrastructure & Site Reliability Engineering - **Company:** Crum & Forster Insurance - **Location:** Glastonbury, CT, United States (Remote available) - **Experience:** Expert - **Salary:** $105,400.0 - $198,100.0 - **Contract:** Permanent contract - **Skills:** Microsoft Access, Artificial Intelligence, Application Performance Management, Microsoft Azure, Bash Shell, Software as a Service, Cloud Computing, Cloud Computing Security, Continuous Integration, DevOps, Identity and Access Management, Information Security Management, Python (Programming Language), Key Management, Network Security, Windows PowerShell, Reliability Engineering, Prometheus, Software Vulnerability Management, Alwayson, Datadog, Grafana, Microsoft Fabric, Information Technology, Bicep, Graphql, Cloud Optimization, Restful APIs, Terraform, Dynatrace, Devsecops, Pagerduty - **Published:** September 10, 2026 - **Apply:** https://www.careerjet.com/job/us339de0a000a2e80d61ef8d9291f89787/eaa ## About the Role * Excellent problem-solving and analytical skills with attention to detail * Strong analytical and problem-solving skills. * Excellent verbal and written communication skills, with the ability to explain technical and functional issues clearly to both technical and non-technical stakeholders. * Demonstrated leadership or mentoring of distributed/offshore teams (formal people-management experience a plus). * Excellent collaboration and influencing skills * Outcome & Metrics Orientation * Self-starter Requirements: * Bachelor's degree in Computer Science, or related field-or equivalent experience. * 8+ years in infrastructure, site reliability, platform, or DevOps engineering, with recent hands-on delivery. * Deep, hands-on expertise operating production workloads on Microsoft Azure. * Proven experience with Infrastructure as Code using Bicep and/or Terraform. * Hands-on experience with observability tooling across metrics, logs, and traces (e.g., Grafana, Prometheus, Azure Application Insights). * Proven experience defining and operating SLIs/SLOs and error budgets. * Hands-on experience securing Azure cloud environments-identity/access, network security, posture management (e.g., Microsoft Defender for Cloud, Azure Policy), and vulnerability management. * Experience scaling Microsoft Fabric and Microsoft Purview. * Experience owning incident management, on-call, and blameless post-incident reviews (e.g., Better Stack, PagerDuty, or comparable). * Strong scripting/automation ability (e.g., PowerShell, Python, Bash) and CI/CD experience (Azure DevOps preferred). * Experience supporting business continuity and disaster recovery with defined RTO/RPO. * Experience building internal developer platforms, golden paths, and self-service infrastructure. * Cloud cost optimization / FinOps experience. * Exposure to AI-assisted operations (AIOps) and modern reliability automation. * Microsoft Azure certifications (e.g., Azure Solutions Architect Expert, Azure DevOps Engineer Expert, Azure Security Engineer Associate) preferred. * Experience in insurance, travel, fintech, or SaaS (regulated environments) preferred. * Prior leadership or mentoring of distributed/offshore teams desired. Formal people-leadership experience a plus ## Description This is a hands-on engineering role with strong leadership influence, focused on reliability and platform. You will lead through technical credibility and influence-setting standards, shaping direction, and elevating the reliability capability across engineering-while remaining deeply hands-on. As TII's Lead Infrastructure & Site Reliability Engineer, you will own the run-time reliability, observability, cloud security, and platform engineering that keep our all-Azure environment secure, resilient, and always-on-supporting the platform and business. This is a build-and-enhance role: you will mature our observability across metrics, logs, and traces; establish disciplined incident command and blameless post-incident practices; codify infrastructure with Bicep and Terraform; harden the security posture of our Azure platform; and build the self-service platform capabilities that let engineering teams move fast safely. You will own reliability, security, and platform hands-on while partnering with engineering architecture, engineering leadership, and security stakeholders. With the AVP, Infrastructure, co-create the future-state vision and roadmap for Cloud, Observability, ITSM, and related domains making it a shared, measurable engineering discipline in support of TII's growth target and expansion into new distribution channels. What you will do: Infrastructure Strategy * Own and evolve TII's Azure infrastructure and reliability roadmap, aligned to the Azure Well-Architected Framework. * With the AVP, Infrastructure to co-create the future-state vision and roadmap for Cloud, Observability, ITSM, and related domains. * Define standards for compute, network, and platform services that scale with business growth and new channels. * Drive cloud cost optimization (FinOps)-balancing performance, resilience, and spend. * Partner with engineering architecture to ensure infrastructure enables design-time resilience and delivery velocity. Site Reliability Engineering * Own run-time reliability across availability, performance, scalability, and capacity for TII's platform. * Mature and expand the SLO practice-defining SLIs, refining the 99.9% (and higher, where warranted) SLOs, and operating error budgets to balance reliability and delivery speed. * Lead capacity planning and performance engineering to support the platform's growth. * Drive operational readiness reviews for new services and major releases. Observability * Own and mature the observability platform across the three pillars-metrics, logs, and traces-enhancing Grafana/Prometheus and Azure Application Insights. * Implement distributed tracing across the GraphQL/REST services to accelerate diagnosis and reduce time-to-detect and time-to-resolve. * Establish meaningful alerting and telemetry that reduce noise and surface real signals. * Build reliability dashboards that give teams and leadership clear visibility into service health. Platform Engineering * Own the internal developer platform and self-service infrastructure capabilities that enable engineering teams to provision and operate safely. * Define golden paths / paved-road templates that make the reliable, secure, and compliant way the easy way. * Establish and champion Infrastructure as Code standards using Bicep and Terraform. * Improve developer experience and engineering enablement through automation and reusable platform services. Operational Excellence & Automation * Drive automation across provisioning, configuration, deployment, and remediation to eliminate toil. * Establish operational runbooks, self-healing patterns, and proactive reliability practices. * Continuously improve deployment safety and rollback capability in partnership with CI/CD owners. Cloud Security & Compliance * Own the engineering and operational security of the Azure cloud platform-including identity and access management, network security, configuration hardening, and secrets/key management. * Manage and improve the cloud security posture (e.g., Microsoft Defender for Cloud, Azure Policy), including continuous vulnerability management and remediation. * Implement security monitoring and alerting as part of the observability platform to detect and respond to threats. * Embed DevSecOps and secure-by-design practices into platform and IaC workflows, enforcing guardrails and policy-as-code within golden paths and self-service tooling. * Partner with the Security function on policy, governance, and compliance in a regulated insurance (PII) environment. Incident & Problem Management * Own the major-incident process and incident command, matured on the on-call platform (Better Stack). * Lead blameless post-incident reviews and drive systemic problem management to prevent recurrence. * Improve on-call health, escalation paths, and mean-time-to-detect / mean-time-to-resolve. * Maintain and enhance the mature Business Continuity and Disaster Recovery capabilities, including RTO/RPO targets and periodic testing. Leadership * Will co-lead a small team of infrastructure, cloud, & system engineers. * Uplift the reliability and platform capability across engineering-raising standards and building a reliability culture. * Mentor engineers, demonstrating the leadership behaviors that support growth into a formal infrastructure leadership role. * Establish standards, documentation, and ways of working that scale across teams. * Other duties as assigned ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Back(end) to the Future: Embracing the continuous Evolution of Infrastructure and Code](https://www.wearedevelopers.com/videos/440-back-end-to-the-future-embracing-the-continuous-evolution-of-infrastructure-and-code) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [The journey from developer to devops - what i've learnt along the way](https://www.wearedevelopers.com/videos/238-the-journey-from-developer-to-devops-what-i-ve-learnt-along-the-way) - [Azure-Well Architected Framework - designing mission critical workloads in practice](https://www.wearedevelopers.com/videos/1529-azure-well-architected-framework-designing-mission-critical-workloads-in-practice) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best Companies to work for in London: Top 25 Companies in 2023](https://www.wearedevelopers.com/magazine/187-best-companies-to-work-for-in-london-top-25-companies-in-2023) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [The 12 Best Jobs for Software Engineers](https://www.wearedevelopers.com/magazine/401-the-12-best-jobs-for-software-engineers)