World Congress 2026 Europe • Jul 9, 2026 • Session details

The Hidden Costs of CPU Limits in Kubernetes

Pavel Malyarevsky

Are your Kubernetes CPU limits secretly spiking tail latency? Unmask the hidden micro-throttling cycle and learn how to instantly eliminate waste by tuning runtime concurrency.

Pause
Mute Enter Fullscreen
#1 about 3 min

Driving factors for CPU limits in shared Kubernetes platforms

Shared Kubernetes platforms use limits to avoid noisy neighbors, ensure policy compliance, and manage multi-tenant budgets.

#2 about 2 min

Virtual machine core allocation versus Kubernetes cgroup throttling

Unlike dedicated virtual machine cores, Kubernetes requests and limits use cgroups that halt workloads when CPU quotas deplete.

#3 about 2 min

Impact of Linux quota replenishment on application latency

Depleted CPU quotas pause workloads until the next 100-millisecond replenishment cycle, increasing latency significantly.

#4 about 3 min

Why CPU usage metrics obscure underlying container throttling

Standard CPU usage metrics fail to indicate throttling because halted containers stop consuming cycles during quota waits.

#5 about 4 min

Demonstrating thread contention and container restarts with node exporter

Testing node exporter with concurrent requests reveals how liveness probe timeouts trigger container restarts during CPU throttling.

#6 about 2 min

Comparing performance improvements by increasing CPU limits

Increasing CPU limits reduces throttling and request failures but fails to address underlying excessive thread concurrency issues.

#7 about 3 min

Optimizing execution time and throttling by reducing concurrent threads

Aligning application thread counts to actual CPU availability drastically reduces throttling and lowers overall CPU consumption.

#8 about 2 min

How language runtimes influence runnable thread concurrency

Different application runtimes inherently produce varying numbers of concurrent threads that exacerbate CPU throttling under limits.

#9 about 3 min

Configuring Go runtime concurrency with the GOMAXPROCS variable

Limiting concurrent threads via runtime settings directly decreases the speed of quota consumption and operating system context switching.

#10 about 4 min

Strategies for balancing hard isolation and application performance

Platform engineers should leverage CPU limits for isolation while tuning thread concurrency based on specific workload service level objectives.

#11 about 2 min

Key takeaways on aligning runtime concurrency with CPU quotas

Matching internal application concurrency with CPU limits is essential for both custom workloads and default platform operators.

#12 about 4 min

Cascading failures and node degradation from unmanaged threads

Unrestricted application threading splits CPU time into impractically small increments and triggers self-amplifying request queues that crash entire nodes.

Matching moments

48 sec

Key takeaways on Go container performance optimization

Rick Rackow Rick Rackow · World Congress 2026 Europe

3:57 min

Common Kubernetes misconfigurations in production environments

Noaa Barki · World Congress 2022

3:13 min

How GOMAXPROCS behaves inside limited container computing environments

Rick Rackow Rick Rackow · World Congress 2026 Europe

2:05 min

Identifying Go runtime overhead in high-throughput systems

Ivan Sinitsin Ivan Sinitsin · Europe 2026 Virtual

1:10 min

Optimizing Kubernetes clusters for resource and cost efficiency

Christian Grieger Christian Grieger · Europe 2026 Virtual

3:05 min

Addressing scale complexities in Kubernetes environments

Bassem Dghaidi · LIVE