World Congress 2026 Europe Jul 9, 2026 Session details

The Hidden Costs of CPU Limits in Kubernetes

Pavel Malyarevsky

Are your Kubernetes CPU limits secretly spiking tail latency? Unmask the hidden micro-throttling cycle and learn how to instantly eliminate waste by tuning runtime concurrency.

Pause
Mute Enter Fullscreen
#1 about 3 min

Driving factors for CPU limits in shared Kubernetes platforms

Shared Kubernetes platforms use limits to avoid noisy neighbors, ensure policy compliance, and manage multi-tenant budgets.

#2 about 2 min

Virtual machine core allocation versus Kubernetes cgroup throttling

Unlike dedicated virtual machine cores, Kubernetes requests and limits use cgroups that halt workloads when CPU quotas deplete.

#3 about 2 min

Impact of Linux quota replenishment on application latency

Depleted CPU quotas pause workloads until the next 100-millisecond replenishment cycle, increasing latency significantly.

#4 about 3 min

Why CPU usage metrics obscure underlying container throttling

Standard CPU usage metrics fail to indicate throttling because halted containers stop consuming cycles during quota waits.

#5 about 4 min

Demonstrating thread contention and container restarts with node exporter

Testing node exporter with concurrent requests reveals how liveness probe timeouts trigger container restarts during CPU throttling.

#6 about 2 min

Comparing performance improvements by increasing CPU limits

Increasing CPU limits reduces throttling and request failures but fails to address underlying excessive thread concurrency issues.

#7 about 3 min

Optimizing execution time and throttling by reducing concurrent threads

Aligning application thread counts to actual CPU availability drastically reduces throttling and lowers overall CPU consumption.

#8 about 2 min

How language runtimes influence runnable thread concurrency

Different application runtimes inherently produce varying numbers of concurrent threads that exacerbate CPU throttling under limits.

#9 about 3 min

Configuring Go runtime concurrency with the GOMAXPROCS variable

Limiting concurrent threads via runtime settings directly decreases the speed of quota consumption and operating system context switching.

#10 about 4 min

Strategies for balancing hard isolation and application performance

Platform engineers should leverage CPU limits for isolation while tuning thread concurrency based on specific workload service level objectives.

#11 about 2 min

Key takeaways on aligning runtime concurrency with CPU quotas

Matching internal application concurrency with CPU limits is essential for both custom workloads and default platform operators.

#12 about 4 min

Cascading failures and node degradation from unmanaged threads

Unrestricted application threading splits CPU time into impractically small increments and triggers self-amplifying request queues that crash entire nodes.

Matching moments

48 sec

Key takeaways on Go container performance optimization

Rick Rackow Rick Rackow · WWC Europe 2026

3:57 min

Common Kubernetes misconfigurations in production environments

Noaa Barki · WWC 2022

3:13 min

How GOMAXPROCS behaves inside limited container computing environments

Rick Rackow Rick Rackow · WWC Europe 2026

3:05 min

Addressing scale complexities in Kubernetes environments

Bassem Dghaidi · LIVE

11:32 min

Audience questions on security, limitations, and Kubernetes crossover

Maurice Brinkmann · LIVE

1:52 min

Utilizing Kubernetes as a foundation for internal platforms

Adam Bien · WWC 2021

Upcoming sessions on this topic

Open session

World Congress 2026 North America

public void saveMoney(AI): The Developer's Guide to Unit Economics

Hrushikesh Pokala

Senior Software Engineer Lead at Equifax

Hrushikesh Pokala
Open session

World Congress 2026 North America

Run your agents in Kubernetes: Build once, deploy anywhere. But really?

Michal Salanci

Senior Systems Engineer at ESET Cybersecurity

Michal Salanci
Open session

World Congress 2026 North America

Trust, But Verify: Continuous GPU Validation at Scale

Kyle Bell

VP of AI @ TensorWave

Kyle Bell
Open session

World Congress 2026 North America

Stop Running Mystery Meat in Production

Jeroen van Erp

Technical Advocate @ SUSE

Jeroen van Erp
Open session

World Congress 2026 North America

How to generate business value through performance optimizations

Nikolai Sidiropulo

Software Engineer at Meta

Nikolai Sidiropulo
Open session

World Congress 2026 North America

KV Cache Is Not About Speed: It's About Surviving Inference Costs

David vonThenen

AI/ML Leader | Keynote Speaker | OSS Engineer & Developer Advocate | Agentic AI, Deep Learning, Production AI | Python, Go, C++

David vonThenen