Journey-Centric Lead Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+9 more
Job description
We are seeking a highly experienced Journey-Centric Lead Site Reliability Engineer (SRE) with Banking and Financial Services (BFS) domain experience to drive end-to-end reliability, observability, automation, and operational excellence across critical customer and business journeys. This role will lead the design and implementation of modern SRE practices, unified observability platforms, self-healing capabilities, AI-driven operations, and workflow automation to ensure highly resilient, scalable, and intelligent digital services., SRE & Reliability Engineering
· Define and implement enterprise-scale SRE best practices across critical applications and digital journeys.
· Establish reliability frameworks, operational standards, and governance models.
· Drive proactive reliability engineering initiatives to improve system availability, resilience, and performance.
· Lead incident management, postmortem analysis, root cause investigations, and reliability reviews.
Unified Observability & Monitoring
· Design and implement a unified observability strategy encompassing metrics, logs, traces, events, and user experience telemetry.
· Build comprehensive observability dashboards for business and technology stakeholders.
· Implement distributed tracing and end-to-end monitoring across complex microservices ecosystems.
· Define observability standards and instrumentation frameworks across engineering teams.
Log Analytics & Trace Correlation
· Enable unified logs, metrics, and trace correlation capabilities for rapid issue detection and troubleshooting.
· Deploy intelligent correlation engines for root cause analysis.
· Improve Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) through observability-driven insights.
· Establish service dependency mapping and journey-centric operational visibility.
Self-Healing & Autonomous Operations
· Design and implement self-healing capabilities using event-driven automation and AI-assisted remediation.
· Develop automated recovery processes for common failure scenarios.
· Create autonomous operational workflows that minimize manual intervention.
· Integrate predictive alerting and automated response mechanisms.
Automation & Workflow Engineering
· Build scalable operational automation frameworks.
· Develop infrastructure, application, observability, and operational workflows using Infrastructure as Code (IaC), Monitoring as Code (MaC), and Observability as Code (OaC).
· Automate deployments, monitoring, remediation, and operational runbooks.
· Reduce operational toil through intelligent engineering solutions.
SLO, SLA & Error Budget Management
· Define and govern measurable Service Level Objectives (SLOs), Service Level Agreements (SLAs), and Error Budgets.
· Partner with engineering and business teams to align reliability targets with customer expectations.
· Establish service maturity metrics and reliability scorecards.
· Drive data-driven operational decision-making through reliability KPIs.
Network & Platform Reliability
· Apply deep understanding of:
o Understanding of network layer to troubleshoot critical bandwidth/latency issues
o No need to pass all these:
o TCP/IP
o DNS
o Load Balancing
o CDN
o API Gateway architectures
o Service Mesh technologies
o VPC and cloud networking
· Troubleshoot complex network performance and availability issues.
· Ensure end-to-end reliability across cloud and hybrid environments.
AI Ops & AWS AgentCore
· Implement and operationalize AI Ops platforms and autonomous operations capabilities.
· Leverage AWS AgentCore to build intelligent operational agents for functions like:
o Incident response
o Root cause analysis
o Predictive remediation
o Capacity forecasting
o Automated operational workflows
· Drive adoption of GenAI-powered operational intelligence across the enterprise.
Requirements
The ideal candidate combines deep SRE expertise with strong platform engineering, cloud operations, networking, observability, and AI Ops experience, with a particular focus on AWS AgentCore-powered operational intelligence and autonomous operations.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path
Is Software Engineering Over-Saturated?
Navigating the AI Shift
Dev Digest 120 - Apple and peers