Coffee With Developers Aug 17, 2026

Why Token Reduction Isn’t Cost Reduction - Dave Anderson & Sarel Weinberger, PhD.

Dave Anderson , Sarel Weinberger , Phd

Compressing LLM inputs doesn't save money. It forces models to work harder, driving up compute costs by 50%. Learn smarter resource allocation techniques that actually lower your bills.

Pause
Mute Enter Fullscreen
#1 about 6 min

The hidden costs of naive token compression

Compressing tokens upfront often forces AI models to do more reading and turns, driving up total computation costs.

#2 about 4 min

Establishing visibility for engineering team budget control

Engineering teams require full end-to-end telemetry to understand the financial impact and ROI of coding agents.

#3 about 6 min

Why most AI context remains unreachable for compression

System prompts, tool definitions, and internal reasoning occupy the vast majority of tokens that naive algorithms cannot optimize.

#4 about 3 min

Creating realistic benchmarks for standard coding agents

Standard benchmarks like SWE-bench misrepresent everyday coding tasks and skew the evaluation of model performance.

#5 about 8 min

How over-optimized context triggers excessive model computation

Removing seemingly irrelevant code causes models to lose confidence and execute expensive re-reads to verify context.

#6 about 8 min

Lowering costs by adjusting harness and effort levels

Utilizing cheaper execution loops and restricting loaded tool definitions drastically reduces unnecessary token generation.

#7 about 5 min

Combating deployment FOMO with practical cost tracking

Implementing a control plane helps teams avoid overspending on high-end models for trivial engineering tasks.

#8 about 6 min

Best practices for interacting with AI coding assistants

Avoid interrupting native harness loops and isolate tasks in fresh sessions to prevent exponential caching costs.

Matching moments

9:14 min

Assessing the token economy and costs of artificial intelligence

Chris Heilmann +2 · LIVE

2:09 min

Shifting focus from token maxing to AI cost management

Dona Sarkar +1 · Coffee With Developers

2:01 min

Transitioning toward AI-first coding and managing token costs

Clemens Wasner Clemens Wasner +4 · World Congress 2026 Europe

1:41 min

Calculating the hidden financial costs of AI coding assistants

Chris Heilmann +2 · LIVE

3:45 min

Optimizing token costs and intelligent chunking strategies

Luca Bianchi Luca Bianchi · World Congress 2026 Europe

3:21 min

Optimizing AI token consumption and running local language models

Chris Heilmann +2 · LIVE

Upcoming sessions on this topic

Open session

World Congress 2026 North America

September 24, 2026 · 14:10–14:40

Stage 1

Anatomy of an AI Request: Where Latency and Cost Are Really Born

Dan Fu

VP of Kernels at Together AI

Dan Fu
Open session

World Congress 2026 North America

September 25, 2026 · 14:10–14:40

Stage 4

Headroom: A Context Optimization Layer for LLM Applications

Tejas Chopra

Senior Software Engineer at Netflix

Tejas Chopra
Open session

World Congress 2026 North America

September 24, 2026 · 11:40–12:10

Stage 6

AI ROI: The Hard Unit Economics of AI-Native Engineering

Manu Gurudatha

Manu Gurudatha, VP of Engineering at PagerDuty

Manu Gurudatha
Open session

World Congress 2026 North America

September 24, 2026 · 12:15–12:45

Stage 4

It's Not About the Models

Bob Wambach

Vice President, Market and Customer Insights for Dynatrace

Bob Wambach
Open session

World Congress 2026 North America

September 25, 2026 · 16:00–16:10

Outdoor Stage

public void saveMoney(AI): The Developer's Guide to Unit Economics

Hrushikesh Pokala

Senior Software Engineer Lead at Equifax

Hrushikesh Pokala
Open session

World Congress 2026 North America

September 24, 2026 · 16:10–16:40

Stage 3

Sandboxing the Swarm: Building Secure, Serverless AI Agents with Wasm

Lena Hall, Thorsten Hans

Lena Hall
Thorsten Hans