> Markdown version of [/videos/1599-keycloak-case-study-making-users-happy-with-service-level-indicators-and-observability?t=141](https://www.wearedevelopers.com/videos/1599-keycloak-case-study-making-users-happy-with-service-level-indicators-and-observability?t=141). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Keycloak case study: Making users happy with service level indicators and observability Are delayed Keycloak logins stalling your workflows? Discover how linking OpenTelemetry traces with SLI metrics pinpoints latency instantly and drastically reduces your mean time to resolution. - **Speakers:** [Alexander Schwartz](https://www.wearedevelopers.com/@alexander-schwartz) - **Event:** World Congress 2025 - **Published:** August 20, 2025 - **Duration:** 27:15 - **URL:** https://www.wearedevelopers.com/videos/1599-keycloak-case-study-making-users-happy-with-service-level-indicators-and-observability ## Summary Modern enterprises depend heavily on Single Sign-On (SSO) solutions like Keycloak, where delayed logins or system downtime drastically degrade user experience and block cross-organization workflows. To guarantee fast and reliable identity services, organizations must actively measure Service Level Indicators (SLIs) to protect against traffic spikes and infrastructure degradation. By implementing robust observability pipelines with tools like Quarkus, Prometheus, and Grafana, DevOps teams can continuously track system availability, pinpoint latency through detailed response-time histograms, and maintain strict Service Level Objectives (SLOs). The real breakthrough in mean time to resolution comes from linking distributed OpenTelemetry tracing with aggregated metrics using exemplars. Exemplars allow engineers investigating a dashboard to click directly from a slow request bucket on a heatmap into the precise Jaeger trace and underlying application log sequence. Capturing just 1% of production traffic, combined with rich business context like client IDs and realm names, is enough to instantly identify failing LDAP systems or exhausted database connection pools. Ultimately, tracking user-centric performance indicators reduces troubleshooting friction and gives engineering leaders concrete data to balance infrastructure optimization with new feature development. **Keywords:** keycloak observability, single sign-on performance, service level indicators, slo monitoring, quarkus metrics, prometheus promql, grafana heatmaps, latency histograms, opentelemetry tracing, trace sampling, metrics exemplars, root cause analysis, identity and access management, jaeger tracing, database connection pooling ## Chapters 1. **Addressing single sign-on challenges in enterprise environments** (00:04) — Why single sign-on stability matters during usage spikes and how observability ensures resilience. 1. **Exploring the Keycloak open-source identity and access system** (02:21) — The core capabilities, licensing, and cloud-native architecture underpinning the Keycloak identity project. 1. **Navigating the primary user and administrator interface screens** (04:05) — How administrators and end-users securely interface with Keycloak login and management consoles. 1. **Establishing core metrics for system reliability and speed** (05:04) — Identifying the critical telemetry data required to maintain fast and dependable authentication protocols. 1. **Defining service level objectives for baseline system performance** (07:21) — Setting quantifiable thresholds for system uptime, connection failures, and overall authentication latency. 1. **Measuring system availability utilizing Prometheus and straightforward PromQL** (09:44) — Calculating overall application uptime using straightforward Prometheus scraping rules and interval queries. 1. **Tracking authentication error rates through basic HTTP metrics** (10:49) — Analyzing web server request counters to consistently evaluate successful versus failed access attempts. 1. **Evaluating request latency using metric distribution timing histograms** (11:31) — Grouping response durations into specific latency buckets to track compliance with acceptable response times. 1. **Visualizing Keycloak performance via standard Grafana troubleshooting dashboards** (13:04) — Monitoring backend system health alongside user response rates inside a centralized visual hub. 1. **Investigating request root causes with distributed tracing tools** (14:08) — Adopting OpenTelemetry tracing frameworks to capture and scrutinize detailed context surrounding slow requests. 1. **Injecting business domain context into distributed authentication traces** (18:00) — Enhancing trace structures with actionable custom identifiers like database operations and token sessions. 1. **Bridging performance metrics and recorded traces using exemplars** (20:19) — Connecting aggregated timing buckets directly to isolated trace records for accelerated anomaly debugging. 1. **Interrogating exemplars via dynamic graphical performance heat maps** (22:15) — Launching targeted diagnostic pathways in Jaeger or Tempo from specific graphical anomaly markers. 1. **Balancing infrastructure optimizations alongside ongoing new feature development** (24:38) — Leveraging centralized operational data to prioritize vital system scaling upgrades over extraneous features. ## Related Moments - [Introduction to metrics and observability challenges in monitoring](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) (from "All your telemetry data from any source in one place") - [Reviewing identity endpoints alongside specific Keycloak preview features](https://www.wearedevelopers.com/videos/1558-delegating-the-chores-of-authenticating-users-to-keycloak) (from "Delegating the chores of authenticating users to Keycloak") - [Utilizing pre-integrated observability and authentication platform tools](https://www.wearedevelopers.com/videos/36-devsecops-security-in-devops) (from "DevSecOps: Security in DevOps") - [Exploring observability dashboards and alert management interface workflows](https://www.wearedevelopers.com/videos/2114-easy-mode-monitoring-and-logging-with-shiftmon) (from "Easy Mode Monitoring and Logging with Shiftmon") - [Defining observability beyond basic metrics, logs, and traces](https://www.wearedevelopers.com/videos/2118-better-together-leveraging-your-observability-tools-as-a-siem) (from "Better Together: Leveraging Your Observability Tools as a SIEM") - [Final performance results and key architectural takeaways](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) (from "Accelerating Authentication Architecture: Taking Passwordless to the Next Level") ## Related Articles - [Events like RSAC Get You CISOs. Developers Decide What Actually Gets Deployed.](https://www.wearedevelopers.com/magazine/693-events-like-rsac-get-you-cisos-developers-decide-what-actually-gets-deployed) - [Devs vs. Marketers, COBOL and Copilot, Make Live Coding Easy and more - The Best of LIVE 2025 - Part 3](https://www.wearedevelopers.com/magazine/663-devs-vs-marketers-cobol-and-copilot-make-live-coding-easy-and-more-the-best-of-live-2025-part-3) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Slopquatting, API Keys, Fun with Fonts, Recruiters vs AI and more - The Best of LIVE 2025 - Part 2](https://www.wearedevelopers.com/magazine/662-slopquatting-api-keys-fun-with-fonts-recruiters-vs-ai-and-more-the-best-of-live-2025-part-2) ## Related Jobs - [Principal Software Engineer, Identity](https://www.wearedevelopers.com/jobs/ext/1469181-principal-software-engineer-identity) at **GitHub** - [Platform Engineer - Mercury Runtime Platform](https://www.wearedevelopers.com/jobs/ext/293235-platform-engineer-mercury-runtime-platform) at **Raiffeisen Bank International AG** - [Senior Engineer, Infrastructure Platform](https://www.wearedevelopers.com/jobs/ext/328836-senior-engineer-infrastructure-platform) at **Intercom, Inc.** - [Senior Backend Engineer (Java)](https://www.wearedevelopers.com/jobs/ext/19369-senior-backend-engineer-java) at **Bonial International GmbH** - [Lead Software Engineer - Data Engineering](https://www.wearedevelopers.com/jobs/ext/2000968-lead-software-engineer-data-engineering) at **Dynatrace** - [Senior Threat Intelligence Analyst](https://www.wearedevelopers.com/jobs/ext/1684162-senior-threat-intelligence-analyst) at **ZEISS Group**