The Cobra Effect : What 150 Years of Bad Incentives Teach Us About AI Token Pricing
- Mahendra Rathod
- 1 day ago
- 14 min read

The Cobra Effect: What Bad Incentives Keep Teaching Us, From Colonial Bounties to AI Token Pricing
In the 1800s, Delhi had a cobra problem. Venomous snakes were common on the streets. People were getting bitten. The British administration running the city wanted them gone.
Their solution was simple. Pay a reward for every dead cobra handed in.
At first it worked. People killed cobras, brought in the bodies, collected the money. Sightings dropped.
Then the numbers started climbing again, month after month. When officials looked into why, they found the answer. People had started breeding cobras. Small pens, regular feeding, a steady supply to sell back to the government. Why hunt a wild, venomous snake when you can farm the reward instead?
When the British worked out what was happening, they cancelled the bounty overnight. The breeders, sitting on pens full of snakes now worth nothing, did the only thing that made sense. They let them go.
Delhi ended up with more cobras than when the scheme started.
This story is famous. It is also probably not true, at least not exactly like this. Historians have gone looking for the breeding-and-release part and found no record of it. It may be a parable that got attached to a real, smaller bounty programme, and grew in the retelling.
But the same thing did happen, for real, somewhere else. Hanoi, 1902, under French colonial rule. The French had built a modern sewer system, and it turned into the best rat habitat the city had ever had. So they paid a bounty, one cent per rat tail. Within months, people were farming rats specifically for their tails. Sewer inspectors started finding tailless rats running around the city — someone had figured out you could cut the tail off a live rat and let it go, so it could grow a new one to harvest later. When the French found actual rat farms outside the city and shut the programme down, the farmers released their rats. This one is documented properly. A historian found the French government's own records in an archive in France and published the whole thing in 2003.
It happened again with fossils. Nineteenth century paleontologists in China and Java paid local diggers by the fragment for dinosaur and hominin bones. The diggers responded exactly as you'd expect. They smashed whole skeletons into as many pieces as possible, because more pieces meant more pay. Years of real scientific evidence got broken to bits for a bounty.
It happened in America too. When the US government paid the builders of the first transcontinental railroad by the mile of track laid, one contractor built an unnecessary curve into the route, adding extra miles just to collect more money.
Four different countries. Four different rewards. Same mistake every time.
You become what you measure.
A hundred and fifty years later, a group of companies in California ran the same experiment. Instead of tails or bones or miles of track, they counted tokens.

Why I'm Writing This
I use paid versions of most major AI tools. Whenever I ask a complex question for research or writing, I hit a limit, and the app pushes me to upgrade. It feels like a free mobile game asking you to buy more lives to keep playing.
My CTO was proud to tell me our whole tech team was using AI to write code, on the premium plans. The real bottleneck, it turned out, was tokens. We managed to stay within the standard plans without paying for extra capacity. Not everyone does.
And every AI company's CEO is on X telling the world how AI is about to change everything. Somewhere underneath that message is a simpler one: buy more tokens.
Those three things made me want to actually work out what a token costs, what we're charged for it, and whether the maths behind AI spending holds up.
Why Companies Blamed Layoffs on AI
In 2023, Klarna's CEO said AI had replaced 700 customer service jobs. It made headlines.
By 2025, Klarna was quietly hiring people back. The AI handled simple questions fine. It struggled with anything that needed judgement — a disputed charge, an angry customer, a complaint that didn't fit the script.
Klarna wasn't unusual. It just got there first.
In April and May of 2026, Meta cut 8,000 jobs. Amazon cut around 30,000. Oracle cut up to 30,000, roughly a fifth of its entire workforce. Microsoft offered voluntary retirement to nearly 9,000 US staff. Snap cut 16% of its workforce. Salesforce's CEO didn't dress it up: "I need less heads," he said, cutting 4,000 customer support roles.
All of it happened while these same companies kept spending enormous amounts on AI. Amazon cut jobs the same quarter AWS grew at its fastest pace in over three years. Alphabet, Microsoft, Meta and Amazon are expected to spend nearly $700 billion combined on AI infrastructure this year, even as they cut headcount.
Marc Andreessen, a venture capitalist with no reason to defend AI hype, said it plainly in an interview: companies over-hired by 25 to 75% during the pandemic, when interest rates were at zero and hiring discipline disappeared. Then rates went up, budgets got tight, and AI became the excuse everyone was waiting for. Now they all have the silver bullet excuse," he said.
Here's the number that gives the game away. At the same time as record layoffs in early 2026, there were 275,000 open AI-related job postings in the US. Hiring for AI roles was up 92% for the year, with a 56% pay premium. The people being let go — customer support, quality assurance, content moderation, middle management — were not the people being hired. Bloomberg's own reporting estimates that roughly half of all AI-attributed layoffs will end with the same roles rehired offshore, or at a lower salary.
A separate Gartner survey found that only 20% of customer service leaders who cut staff did so because AI had actually replaced the work. The rest were responding to ordinary economic pressure and used AI as convenient cover.
Many of these companies had already decided to cut costs. AI gave them a story that sounded better than the real one.

The Cost Equation Nobody Is Actually Calculating
Whatever the real reason for the layoffs, the bet has been placed. Fewer people, more AI. There's only one honest way to know if that bet is working, and almost nobody has actually done it.
OLD WORLD COST = Full Team Salaries + Overhead
NEW WORLD COST = Retained Team Salaries + Overhead
+ AI Subscriptions
+ Token Costs
− Savings From LayoffsFor the new world to be cheaper, the second number has to come out smaller than the first. That's the whole test. And the one part of that equation that changes week to week, that almost no company has actually modelled, is the cost of tokens.
So we need to understand what a token is, first.
What Is a Token, and How Is It Priced
A token is roughly four characters, about three-quarters of a word in English. "The cat sat on the mat" is about seven tokens. This blog post costs around six thousand.
For a single question, that's nothing. It stops being nothing once you look at what a real task costs.
Look at the last row on that table. A single agentic coding task — the kind Cursor or Claude Code run when actually building something — can burn between one and three and a half million tokens. That's because these tools re-read everything they've done so far, every single step, before deciding the next move. The cost keeps stacking.
There's also a cost you never see. Every AI product runs hidden instructions on every message you send — a system prompt, usually running five hundred to three and a half thousand tokens, every single time, before you've typed a word. A company running an internal chatbot for ten thousand messages a day pays that cost up to thirty-five million times over, for nothing the user ever asked for.
And the cost is not symmetric. Output tokens — what the AI writes back — cost four to six times more than input tokens on most frontier models. DeepSeek is the exception, at around two times. When an AI tool writes you three hundred lines of code, that output costs far more than reading your files did. Most people only notice this once the bill arrives.
The real cost is not in the answer you get back. It's in every step the model takes to reach it.

What It Actually Costs to Produce a Token
What you're charged for a token has almost nothing to do with what it costs to make one.
At the raw compute level, a token costs about $0.013 per million to produce. Practically free.
But that's only the electricity bill. The full cost includes the chips, the training runs, the research teams, the safety teams, and spare capacity kept ready for traffic spikes.
Add it all up, and here's what it costs a company like OpenAI to stay in business:
OpenAI lost roughly $5 billion in 2024, against $3.7 billion in revenue. In the first half of 2025, losses ran between $7 and $13 billion. Their own forecast projects a $14 billion loss for 2026. ChatGPT reportedly costs $700,000 a day just to run.
This isn't a startup burning through seed funding while it finds its feet. This is the largest company in the industry losing more than $2 billion a month, on purpose, because charging what tokens actually cost would mean charging you far more than you pay today.
The price on your invoice is not the real price. It's a subsidized price, and subsidies don't last forever.
Why Enterprises Are Getting Angry About AI Pricing
If this subsidy argument sounds like something only economists worry about, it isn't. The people paying these bills are saying so, loudly, in public.
In July 2026, Palantir's CEO Alex Karp went on CNBC and said this, on the record:
"I am paying for tokens that create no value. These people are stealing the weights and alpha of my business, and they're creating a wealth tax that does not help the poor, it just punishes."
Worth being upfront about: Karp has a direct commercial interest here. Palantir sells enterprises a different model — one where the customer keeps their own data and their own model weights, instead of renting access to someone else's. This is a competitor attacking a rival's business model, on television, in front of the exact customers he wants to win.
That doesn't make his underlying question wrong. If a frontier AI model genuinely delivered the value the labs claim, why charge by the token at all? Why not take a share of the value created instead, the way a confident advisor would?
Charging by the token, in this reading, tells you something. If the product reliably created the value being claimed, the seller would price for the value. They price for compute instead, because compute is the only thing they can measure for certain.
Palantir's own public materials on AI sovereignty make a related point worth taking seriously, separate from Karp's television appearance: your data is your compounding advantage, and handing it to a third party risks handing over the edge that makes your business worth running. Your model's weights are the distilled form of everything your business has learned. And every prompt you send is, in effect, a lesson for someone else's model.
Every question you send a model teaches it something about your business. That knowledge does not stay yours.

What You're Being Charged Today
Here's the current pricing, checked against official provider pages as of July 6, 2026:
And here's how that price has moved over time:
Frontier pricing has fallen 83% in three years. GPT-4 launched at $30 per million input tokens in 2023. GPT-5.5 sits at $5 today. Budget-tier models have fallen further still, down to a few cents. Both drops are real, and mostly explained by competition — DeepSeek reportedly trained a frontier-competitive model for around $6 million, against GPT-4's roughly $100 million.
But a lower price doesn't always mean proportionally less intelligence:
DeepSeek's current flagship trails GPT-5.5 and Gemini 3.1 Pro by under seven points on a demanding reasoning benchmark, while costing a fourteenth to a thirty-fifth of the price. As one AI newsletter put it recently: sending every query to the most expensive model is like hiring a surgeon to apply a band-aid.
Why Falling Prices Are Not Lowering Anyone's Bill
Here's the number that should stop every CFO mid-sentence.
Weekly token use across the market grew roughly twelve times over between early 2025 and May 2026. Over the same period, the price per token kept falling.
Cheaper per unit. More expensive overall. This is called the Jevons Paradox, after a 19th century economist who noticed that when coal got cheaper, Britain used more of it, not less, and total spending on coal went up. Cheap power unlocked uses nobody had bothered with at the old price.
The reason, in AI's case, is agentic work. A simple chat exchange might use two thousand tokens. An agentic workflow — one task passed between a main agent and several smaller agents, running tool calls, retrying failed steps — can burn five hundred thousand tokens to finish a job that used to take one exchange.
This is exactly what happened at Uber.
In December 2025, Uber gave 5,000 engineers a new AI coding tool and told them to use it heavily. By February 2026, 32% were using its most advanced features. By March, 84%. In one month.
By April, the budget was gone.
Not slightly over. Gone. A full year's allocation for one tool, burned through in four months, at $500 to $2,000 per engineer per month.
The cause wasn't misuse. It was an internal leaderboard ranking teams by how many tokens they used. Not by what they shipped. By usage.
Uber's engineers weren't behaving badly. They were behaving exactly like the cobra breeders, or the rat farmers in Hanoi, or the fossil diggers in China. Someone set up a reward for the wrong thing, and people responded to the reward, not the intention behind it. Uber's own COO admitted, on the record, that the link between all that AI-written code and anything genuinely useful to customers "is not there yet."
This is the same mistake as the cobra breeders. A leaderboard instead of a bounty. Same result.
Worth noting where this money actually goes. Engineering accounts for more than 60% of enterprise AI spend, at per-person costs over ten times higher than sales teams. When a company asks whether AI is "worth it," the honest answer usually depends on what one department is doing, not a company-wide average.
Does the Cost Equation Actually Work?
Now the maths this whole piece has been building toward.
I'm going to run this for a mid-size software company with 20 developers. India-based team first, then a US-based one. One assumption, stated plainly: AI coding tools deliver somewhere between a 20% and 40% productivity gain. Real studies disagree sharply on this — one controlled study found negative productivity for experienced developers on complex work, another found 55% gains on well-defined tasks. This range is a working assumption for the maths below, not a settled number.
An India-based team of 20, at roughly ₹91.5 lakh a month in fully-loaded cost, needs AI to deliver at least a 17.5% productivity gain before the new setup actually costs less than the old one.
A US-based team, at roughly ₹2.44 crore a month, only needs a 6.6% gain to break even.
The cheaper your people, the higher the bar AI has to clear to beat them on cost. "AI will replace Indian IT" is a far more complicated claim than the headlines suggest, because Indian IT's own low cost is its best defence.
There's a wrinkle worth sitting with here. The subscription you pay isn't subsidised evenly. One developer who tracked their own usage on a $200 Claude Max plan found their actual token use would have cost roughly $3,650 a month at metered API rates — an eighteen-fold discount. Anthropic's own reported average for Claude Code developers sits closer to $6 a day, nowhere near that ceiling.
Subscriptions work like gym memberships. A few people show up every day and use every machine, and for them the flat fee is an enormous bargain. Most people pay the same fee and barely go. The gym survives because of that mismatch, not despite it. AI subscriptions run the same trade. The heaviest users get the wildest discount. Everyone else quietly funds it.
What Token Prices Need to Fall To
Working backwards from the India team's 10% productivity scenario: for the equation to close at that modest a gain, AI tooling cost per developer needs to stay under $480 a month. Subtract the $40 subscription, and that leaves $440 for tokens. At 15 to 20 million tokens a month of moderate use, that implies an average price of $22 to $29 per million tokens.
Current pricing for output-heavy agentic work runs well above that. So at a 10% gain, the India team's equation doesn't close yet on intensive agentic work. At 25 to 30%, it closes comfortably. For the US team, it closes at almost any realistic gain, even at today's prices.
Prices are moving in the right direction. DeepSeek's own flagship price fell 75% in the three months before this was written. Companies that send simple work to cheap models, and save expensive models for genuinely hard problems, will get to a positive equation faster than companies that use the priciest option for everything. I say this as someone whose Claude subscription runs out faster than any of my other three. There's a lesson in that for me too.
Two Different AI Markets Are Forming
Citadel Securities, the trading firm, published a research note in June 2026 worth sitting with here.
Their argument: the real limits on AI were never about capability. They were always going to be physical — compute, power, cooling — and those limits are now showing up. Microsoft cancelling Claude Code subscriptions. Amazon quietly removing its own usage leaderboard. Companies discovering bills far larger than they modelled.
Out of this, Citadel argues, two separate AI markets are forming. Frontier models will concentrate among companies with the balance sheets and the genuinely hard problems that justify the cost. Everyone else moves to cheaper, purpose-built models for routine work. The data already shows this happening — total AI spending across the market has started shrinking for the first time in months, as companies switch to cheaper options rather than giving up on AI altogether.
The choice most companies face isn't whether to use expensive AI. It's knowing which tasks actually need it.

What This All Means
Three patterns show up in everything above.
Some companies used AI as cover for cuts that were already coming. They over-hired through the boom years, came under pressure to cut costs, and AI gave them a story that sounded better than "we hired too many people." Klarna made it public first. Meta, Amazon, Oracle, Salesforce and Microsoft ran the same play in 2026, at a much bigger scale.
Some companies deployed AI without doing this maths first, measured the wrong thing, and got burned. Uber is the clearest case. Rewarding usage instead of output is not a rare mistake. It shows up everywhere once you know to look for it.
And some companies are quietly getting it right. They treat AI as a genuine complement to their people, measure what actually got built rather than what got consumed, and make unglamorous progress that never becomes a headline, because it isn't a story. It's just good management.
Every reward eventually gets tested by what it actually measures. The British paid for dead cobras, or something close to it, and the French definitely paid for rat tails, and got farmed rats. AI labs charge for tokens used, and got Uber's leaderboard.
So here's the question nobody in this industry seems eager to answer: if AI token pricing keeps producing this exact mistake, what would a better model actually look like?
Charge per API call, and someone just games the call count instead. Charge based on outcomes, and you need a definition of "outcome" precise enough to bill against, which nobody has solved yet. Take equity instead of fees, and every AI company quietly turns into a venture fund with better marketing.
None of these are simple fixes. That's the point.
Thirty years of software pricing was built on seats, licenses, and predictable monthly bills. AI has broken that model, and nobody has replaced it with something better yet. Whoever works out the real mechanism won't just win the AI market. They'll change how every software company charges for everything that comes after.
You become what you measure. Nobody in this industry has decided what to measure yet.

Happy Reading!
Further Reading
Antifragile by Nassim Nicholas Taleb
About systems that get stronger from shocks and disorder, instead of just surviving them. The whole idea of measuring the wrong thing and getting punished for it later comes straight from Taleb's thinking. This book is the deeper version of the cobra story.
Prediction Machines by Ajay Agrawal, Joshua Gans, and Avi Goldfarb
Explains AI as a prediction technology, and what that means for where it actually creates value. Useful for understanding why AI works well for coding and content, and struggles almost everywhere else — a question this blog keeps circling back to.
The Second Machine Age by Erik Brynjolfsson and Andrew McAfee
On why the gains from new technology usually take years to show up, even after everyone's already adopted it. Relevant to the companies that fired people expecting instant savings. The book explains why those savings, if they're real at all, might still be a few years away.
Skin in the Game by Nassim Nicholas Taleb
About who actually pays the price when a decision turns out wrong. Every CEO who announced AI layoffs was risking someone else's job, not their own. This book is the sharpest lens I know for that kind of asymmetry.



Comments