Amazon ended up with unexpectedly large charges while using Claude from Anthropic. A seemingly simple programming task cost approx. $1.8M after an internal project ran about 860% over its planned budget. The culprit was a configuration/billing error that went unnoticed for five months — FYI, what teams elsewhere called “extremely cheap” to run turned into what sources labelled “catastrophically expensive.”

When the overspend was discovered, Amazon put in automatic spending caps and extra monitoring, and it shut down an internal leaderboard that rewarded high token use. The episode is a blunt example of how small infra mistakes and misaligned incentives (e.g., chasing token-heavy outputs) can balloon into serious bills once LLMs are used at scale — billing is boring until it isn’t.