> Markdown version of [/jobs/ext/2246564-journey-centric-lead-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2246564-journey-centric-lead-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Journey-Centric Lead Site Reliability Engineer - **Company:** SELF Inc - **Location:** United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Cloud Computing, Complex Networks, Domain Name System (DNS), Monitoring of Systems, Network Layer, Log Analysis, Reliability Engineering, Site Reliability Engineering Practices, TCP/IP, Load Balancing, Istio, System Availability, Mttr, Infrastructure as Code (IaC), Amazon Virtual Private Cloud (VPC), Deployment Automation, Api Gateway, Dynatrace, Microservices - **Published:** August 26, 2026 - **Apply:** https://www.dice.com/job-detail/82266df6-2d17-49f1-94c8-a03eb8ba66a9 ## About the Role The ideal candidate combines deep SRE expertise with strong platform engineering, cloud operations, networking, observability, and AI Ops experience, with a particular focus on AWS AgentCore-powered operational intelligence and autonomous operations. ## Description We are seeking a highly experienced Journey-Centric Lead Site Reliability Engineer (SRE) with Banking and Financial Services (BFS) domain experience to drive end-to-end reliability, observability, automation, and operational excellence across critical customer and business journeys. This role will lead the design and implementation of modern SRE practices, unified observability platforms, self-healing capabilities, AI-driven operations, and workflow automation to ensure highly resilient, scalable, and intelligent digital services., SRE & Reliability Engineering · Define and implement enterprise-scale SRE best practices across critical applications and digital journeys. · Establish reliability frameworks, operational standards, and governance models. · Drive proactive reliability engineering initiatives to improve system availability, resilience, and performance. · Lead incident management, postmortem analysis, root cause investigations, and reliability reviews. Unified Observability & Monitoring · Design and implement a unified observability strategy encompassing metrics, logs, traces, events, and user experience telemetry. · Build comprehensive observability dashboards for business and technology stakeholders. · Implement distributed tracing and end-to-end monitoring across complex microservices ecosystems. · Define observability standards and instrumentation frameworks across engineering teams. Log Analytics & Trace Correlation · Enable unified logs, metrics, and trace correlation capabilities for rapid issue detection and troubleshooting. · Deploy intelligent correlation engines for root cause analysis. · Improve Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) through observability-driven insights. · Establish service dependency mapping and journey-centric operational visibility. Self-Healing & Autonomous Operations · Design and implement self-healing capabilities using event-driven automation and AI-assisted remediation. · Develop automated recovery processes for common failure scenarios. · Create autonomous operational workflows that minimize manual intervention. · Integrate predictive alerting and automated response mechanisms. Automation & Workflow Engineering · Build scalable operational automation frameworks. · Develop infrastructure, application, observability, and operational workflows using Infrastructure as Code (IaC), Monitoring as Code (MaC), and Observability as Code (OaC). · Automate deployments, monitoring, remediation, and operational runbooks. · Reduce operational toil through intelligent engineering solutions. SLO, SLA & Error Budget Management · Define and govern measurable Service Level Objectives (SLOs), Service Level Agreements (SLAs), and Error Budgets. · Partner with engineering and business teams to align reliability targets with customer expectations. · Establish service maturity metrics and reliability scorecards. · Drive data-driven operational decision-making through reliability KPIs. Network & Platform Reliability · Apply deep understanding of: o Understanding of network layer to troubleshoot critical bandwidth/latency issues o No need to pass all these: o TCP/IP o DNS o Load Balancing o CDN o API Gateway architectures o Service Mesh technologies o VPC and cloud networking · Troubleshoot complex network performance and availability issues. · Ensure end-to-end reliability across cloud and hybrid environments. AI Ops & AWS AgentCore · Implement and operationalize AI Ops platforms and autonomous operations capabilities. · Leverage AWS AgentCore to build intelligent operational agents for functions like: o Incident response o Root cause analysis o Predictive remediation o Capacity forecasting o Automated operational workflows · Drive adoption of GenAI-powered operational intelligence across the enterprise. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Turning Container security up to 11 with Capabilities](https://www.wearedevelopers.com/videos/718-turning-container-security-up-to-11-with-capabilities) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship)