Coffee With Developers Aug 17, 2026

Why Token Reduction Isn’t Cost Reduction - Dave Anderson & Sarel Weinberger, PhD.

Dave Anderson , Sarel Weinberger , Phd

Compressing LLM inputs doesn't save money. It forces models to work harder, driving up compute costs by 50%. Learn smarter resource allocation techniques that actually lower your bills.

Pause
Mute Enter Fullscreen
#1 about 6 min

The hidden costs of naive token compression

Compressing tokens upfront often forces AI models to do more reading and turns, driving up total computation costs.

#2 about 4 min

Establishing visibility for engineering team budget control

Engineering teams require full end-to-end telemetry to understand the financial impact and ROI of coding agents.

#3 about 6 min

Why most AI context remains unreachable for compression

System prompts, tool definitions, and internal reasoning occupy the vast majority of tokens that naive algorithms cannot optimize.

#4 about 3 min

Creating realistic benchmarks for standard coding agents

Standard benchmarks like SWE-bench misrepresent everyday coding tasks and skew the evaluation of model performance.

#5 about 8 min

How over-optimized context triggers excessive model computation

Removing seemingly irrelevant code causes models to lose confidence and execute expensive re-reads to verify context.

#6 about 8 min

Lowering costs by adjusting harness and effort levels

Utilizing cheaper execution loops and restricting loaded tool definitions drastically reduces unnecessary token generation.

#7 about 5 min

Combating deployment FOMO with practical cost tracking

Implementing a control plane helps teams avoid overspending on high-end models for trivial engineering tasks.

#8 about 6 min

Best practices for interacting with AI coding assistants

Avoid interrupting native harness loops and isolate tasks in fresh sessions to prevent exponential caching costs.

Matching moments

9:14 min

Assessing the token economy and costs of artificial intelligence

Chris Heilmann +2 · LIVE

2:09 min

Shifting focus from token maxing to AI cost management

Dona Sarkar +1 · Coffee With Developers

2:01 min

Transitioning toward AI-first coding and managing token costs

Clemens Wasner Clemens Wasner +4 · World Congress 2026 Europe

1:41 min

Calculating the hidden financial costs of AI coding assistants

Chris Heilmann +2 · LIVE

3:45 min

Optimizing token costs and intelligent chunking strategies

Luca Bianchi Luca Bianchi · World Congress 2026 Europe

3:21 min

Optimizing AI token consumption and running local language models

Chris Heilmann +2 · LIVE

Upcoming sessions on this topic

Open session

World Congress 2026 North America

Headroom: A Context Optimization Layer for LLM Applications

Tejas Chopra

Senior Software Engineer at Netflix

Tejas Chopra
Open session

World Congress 2026 North America

AI ROI: The Hard Unit Economics of AI-Native Engineering

Manu Gurudatha

Manu Gurudatha, VP of Engineering at PagerDuty

Manu Gurudatha
Open session

World Congress 2026 North America

public void saveMoney(AI): The Developer's Guide to Unit Economics

Hrushikesh Pokala

Senior Software Engineer Lead at Equifax

Hrushikesh Pokala
Open session

World Congress 2026 North America

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
Open session

World Congress 2026 North America

KV Cache Is Not About Speed: It's About Surviving Inference Costs

David vonThenen

AI/ML Leader | Keynote Speaker | OSS Engineer & Developer Advocate | Agentic AI, Deep Learning, Production AI | Python, Go, C++

David vonThenen
Open session

World Congress 2026 North America

The Things Your AI Isn't Telling You

Desmond Lamptey

Lead Software Engineer @ Capital One

Desmond Lamptey