Coffee With Developers
•
Aug 17, 2026
Why Token Reduction Isn’t Cost Reduction - Dave Anderson & Sarel Weinberger, PhD.
Dave Anderson , Sarel Weinberger , Phd
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Compressing LLM inputs doesn't save money. It forces models to work harder, driving up compute costs by 50%. Learn smarter resource allocation techniques that actually lower your bills.
Matching moments
More from Coffee With Developers
Related videos
Related articles
DC
Daniel Cranney
ME
Markus Eisele
DC
Daniel Cranney
DC
Daniel Cranney
CH
Chris Heilmann
From learning to earning
Jobs that call for the skills explored in this talk.
23 days ago
Principal Software Engineer, AI Inference Runtime
ARM
Seattle, WA, United States
Expert
$262k
Compilers
Low Latency
Concurrency
23 days ago
Staff Software Engineer, AI Inference Runtime
ARM
Seattle, WA, United States
Expert
$209k–282k
Compilers
Low Latency
Concurrency
about 1 month ago
•
Verified
LLM Training Engineer
Sciforium
San Francisco, United States
Expert
$155k–220k
Python
23 days ago
Principal Software Engineer, AI Inference Cloud
ARM
Seattle, WA, United States
Expert
$262k
Pytorch
TensorRT
Kubernetes
24 days ago
Senior AI Developer
PwC
United States
Expert
Remote
Docker
Github
Fastapi
about 1 month ago
•
Verified
Partner Sales Director - AI Alliances - Model Providers
Dynatrace
San Francisco, United States
Expert
Remote
DevOps
AI Frameworks
Machine Learning