Skip to content
All articles
AI & Business11 min read

AI Got Cheaper. So Why Are Companies Suddenly Scared of the Bill?

The price of AI tokens keeps falling, yet companies are discovering that their AI bills can still explode. From employees casually burning millions of tokens to autonomous agents running hundreds of model calls behind the scenes, the real AI cost problem is no longer the price of intelligence. It is how much intelligence we are willing to consume.

AI companies have spent the last few years telling us the same beautiful story.

Models are getting smarter.

Models are getting cheaper.

Inference is getting faster.

Soon, intelligence will be almost free.

There is just one problem.

Companies are discovering that when something becomes cheap enough, people start using a ridiculous amount of it.

And that is where the AI bill gets interesting.

A company can save 90% on the cost of a single AI request and still spend more money on AI six months later.

Not because the technology got more expensive.

Because everybody started using it for everything.

That is the part of the AI economy nobody really wanted to talk about.

Until the invoices started arriving.

THE TOKEN BILL NOBODY WAS WATCHING

For ordinary users, AI pricing feels simple.

Pay $20 a month.

Open ChatGPT.

Ask questions.

Generate some images.

Write some emails.

Maybe complain when you hit a usage limit.

Enterprise AI doesn't work like that.

Behind the friendly interface is a meter.

Tokens go in.

Tokens come out.

Someone pays for them.

And the number of tokens an AI system consumes can become enormous once you stop using it like a chatbot and start using it like an employee.

A simple question might require thousands of tokens.

A research task might require tens of thousands.

A coding agent can use hundreds of thousands or millions.

A complex agentic workflow can go even further.

One recent study of eight frontier models found that agentic coding tasks consumed roughly 1,000 times more tokens than ordinary code reasoning or code chat. Even more interestingly, the researchers found that using more tokens did not reliably mean getting better results. Some tasks actually peaked at an intermediate level of spending.

That should make every CFO interested in AI sit up.

Because suddenly the question isn't:

“How much does the model cost?”

It is:

“How much work are we allowing the model to do?”

THE CHEAP AI TRAP

This is where things get slightly counterintuitive.

Imagine an AI provider cuts its price by 90%.

Fantastic.

Your finance department celebrates.

Then the company starts putting the model everywhere.

Customer support gets an AI assistant.

Sales gets an AI assistant.

Marketing gets one.

HR gets one.

Developers get coding agents.

Managers get meeting summarizers.

Legal gets document analysis.

Operations gets background agents.

Someone connects an AI to the company's entire document archive.

Another person builds an automated workflow that wakes up every morning and asks five different models to analyze yesterday's data.

Suddenly the company isn't buying one AI product.

It is operating a small artificial workforce.

And artificial workers do not clock out.

They just consume tokens.

This is one reason enterprise AI spending is becoming a bigger issue even while the raw price of inference continues to fall.

Recent reporting from 404 Media described companies trying to control runaway token consumption, including employees using AI for surprisingly trivial tasks. Accenture, according to leaked audio reported by the publication, was dealing with employees burning through tokens on things as mundane as converting PDFs into presentations.

That sounds funny.

It isn't.

Because the PDF is not the problem.

The problem is that the person pressing the button doesn't see the meter.

THE INVISIBLE METER

Imagine if every time you opened Microsoft Word, your company charged you based on how many characters you typed.

You would notice.

People would become very economical with their words.

Now imagine the meter was completely invisible.

Write a 50-page report?

Fine.

Ask the AI to rewrite it five times?

Fine.

Ask it to summarize the same document again because you didn't like the first summary?

Fine.

Then send the output into another model.

Then another.

Nobody sees the bill until the end of the month.

That's roughly what is happening with AI consumption.

The user sees a chat window.

The company sees infrastructure consumption.

Those are two very different experiences.

And agentic systems make the problem worse.

A human asks an agent to “research this company.”

The agent searches.

Reads documents.

Calls a model.

Calls another tool.

Reads the results.

Reasons about them.

Writes a draft.

Checks the draft.

Searches again.

Rewrites it.

Maybe asks another model to verify it.

The user sees one request.

The infrastructure sees a small army.

That distinction is going to become one of the biggest economic questions in enterprise AI.

THE TOKEN ECONOMY IS GETTING ITS OWN FINANCE DEPARTMENT

Something strange has started happening.

Companies are building financial controls specifically for AI consumption.

The Linux Foundation recently launched an initiative around AI token economics with major companies backing the effort, reflecting a growing need for standardized ways to understand and manage AI usage.

BCG has also described the next stage of enterprise AI governance as moving from controlling access to controlling economics.

That sounds incredibly boring.

It isn't.

Because it means AI is slowly turning into a utility.

Think electricity.

Nobody asks how much electricity a single lightbulb costs.

They ask how much electricity the building consumes.

That's the shift happening with AI.

The individual API request doesn't matter nearly as much as the workload.

And once companies start measuring workloads instead of prompts, some uncomfortable questions appear.

Why are we using the most expensive model for this?

Why are we sending the entire document every time?

Why is this agent making twelve calls when three would work?

Why are we using a frontier reasoning model to classify customer emails?

Why are we paying for a million-token context when the agent only needs ten thousand?

Why does this automated workflow run every five minutes?

Those questions can save far more money than negotiating another 5% discount with the model provider.

THE MODEL IS NOT ALWAYS THE EXPENSIVE PART

Here's the part that gets missed in most AI pricing discussions.

The model isn't necessarily your biggest cost.

The workflow can be.

Suppose Model A costs twice as much as Model B.

That doesn't automatically mean Model B is cheaper.

If Model B takes three attempts to complete something that Model A handles correctly on the first attempt, you might end up spending more.

A recent enterprise coding-agent study found exactly this kind of tradeoff when comparing cloud APIs with on-premise models. The researchers found that aggressive prompt caching could dramatically reduce API costs, while some lower-cost local configurations created substantially more defect-repair work.

That is a much more interesting way to think about AI economics.

Not:

“Which model has the cheapest tokens?”

But:

“What does a successful task cost us?”

That number includes inference.

Retries.

Tool calls.

Human review.

Failed outputs.

Debugging.

Latency.

Infrastructure.

Storage.

Monitoring.

And sometimes the cost of an employee fixing what the AI confidently broke.

THE MOST EXPENSIVE AI MIGHT BE THE ONE THAT FEELS FREE

This is where companies need to be careful with unlimited AI subscriptions.

Unlimited sounds wonderful.

But unlimited access can create unlimited experimentation.

And experimentation is exactly what people do when there is no visible cost.

An employee who would never ask an expensive consultant to rewrite a 60-page document three times might happily ask an AI.

An engineer who would normally spend ten minutes understanding a small problem might throw it at an agent.

A marketer might generate 200 versions of an advertisement instead of five.

A manager might have an AI summarize every meeting, every document and every email.

None of these decisions looks financially significant.

Multiply them across 10,000 employees.

Now they are infrastructure.

This is the AI version of the cloud bill problem.

Cloud computing didn't become expensive because servers suddenly became more expensive.

Companies became good at creating servers.

Then they created too many.

AI may be following the same path.

WE MAY HAVE INVENTED THE AI VERSION OF JEVO​​NS' PARADOX

There is an old economic idea that when technology makes a resource more efficient, people don't necessarily use less of it.

They often use more.

Make computing cheaper.

More computing gets used.

Make storage cheaper.

More data gets stored.

Make bandwidth cheaper.

More video gets streamed.

AI may be heading in the same direction.

Make intelligence cheaper.

People find more things to make intelligent.

That's why falling token prices don't automatically mean falling AI expenditure.

The cheaper the intelligence becomes, the more workflows become economically viable.

And some of those workflows would have looked completely ridiculous two years ago.

Nobody would have hired a human to summarize every internal document every morning.

But if an AI agent can do it for a few dollars?

Why not?

Then someone decides every department needs one.

Then every department needs five.

Then somebody connects them together.

Now you're paying for an AI bureaucracy.

THE MILLION-TOKEN PROBLEM

There is another hidden cost here.

Context.

Every time an AI system processes a large amount of information, somebody is paying for that information to be processed.

This is why enormous context windows are both impressive and dangerous.

A one-million-token context window sounds like unlimited memory.

But if your agent keeps throwing huge amounts of context into every request, you can burn through an enormous amount of money without producing proportionally better results.

Research into token-efficient inference is already looking at ways to route workloads according to their actual token requirements because blindly provisioning for maximum context can waste substantial infrastructure capacity. One recent study found potential GPU savings of 17% to 39% through token-budget-aware routing.

This is where good AI engineering starts looking suspiciously like good old-fashioned systems engineering.

Cache aggressively.

Route intelligently.

Use smaller models when they are good enough.

Batch what doesn't need to be instant.

Keep context focused.

Measure actual workloads.

Stop pretending every request deserves the biggest model.

THE AI MODEL YOU USE MIGHT MATTER LESS THAN HOW YOU USE IT

This is probably the biggest lesson from the current AI spending problem.

Companies have spent enormous amounts of time asking:

Which model is smartest?

Which model writes better code?

Which model has the biggest context?

Which model has the highest benchmark score?

Those are useful questions.

But once AI becomes infrastructure, another question becomes more important:

Which model is appropriate for this job?

You don't use a Formula 1 car to deliver groceries.

You don't need a frontier reasoning model to categorize 10,000 customer emails.

You don't need a million-token context window to answer a question about a five-line database query.

You don't need an autonomous agent to rename a file.

And you probably don't need the most expensive model available to summarize yesterday's meeting.

The smartest model is not always the cheapest solution.

Sometimes the smartest architecture is knowing when not to use it.

AI FINOPS IS ABOUT TO BECOME A REAL JOB

There is a new role hiding underneath all of this.

Someone has to understand:

Which models are being used.

Which teams are consuming them.

What workflows generate the most tokens.

Where retries are happening.

Where caching can help.

Which models are overkill.

Which agents are wasting context.

Which AI features actually generate revenue.

And which ones exist because somebody thought they sounded cool during a strategy meeting.

That person is going to become very important.

Call it AI FinOps.

Call it AI infrastructure economics.

Call it whatever you want.

The job is essentially the same:

Make sure the company gets more useful work out of every dollar spent on inference.

That's already becoming a serious business opportunity. Companies are emerging specifically to help enterprises measure and reduce AI spending, while major consulting and technology firms are building their own AI cost-management offerings.

And honestly, it makes sense.

The AI industry spent years teaching companies how to adopt AI.

Now someone has to teach them how to afford it.

THE REAL AI BILL IS NOT THE API BILL

This is the part I think companies are going to learn the hard way.

The invoice from OpenAI, Anthropic, Google, xAI or whoever is only one line in the spreadsheet.

The real cost of AI is:

Inference

Infrastructure

Integration

Security

Monitoring

Human review

Failures

Retries

Data preparation

Context management

Vendor lock-in

And the organizational cost of changing how people work.

An AI agent that costs $500 a month but saves $10,000 in employee time is cheap.

An AI agent that costs $50 a month but creates $5,000 of cleanup work is expensive.

The token price doesn't tell you which one you bought.

That is why the next phase of AI adoption is going to be less exciting than the last one.

There will be fewer “look what the model can do” demos.

More spreadsheets.

More dashboards.

More usage limits.

More routing.

More caching.

More smaller models.

More arguments between engineering and finance.

And probably a lot more people asking:

“Why did our AI bill jump 400%?”

THE AI GOLD RUSH IS ENTERING THE ACCOUNTING DEPARTMENT

This doesn't mean AI is a bad investment.

Quite the opposite.

It means AI is becoming real enough to have boring problems.

That's actually a good sign.

Every major technology eventually reaches this stage.

First, everyone asks whether it works.

Then everyone wants it.

Then everyone discovers it costs money.

Then finance arrives.

AI is entering that third stage now.

The companies that win won't necessarily be the ones that spend the most on models.

They'll be the ones that figure out where expensive intelligence actually creates value.

They'll use the giant model when the giant model is worth it.

They'll use a small model when it isn't.

They'll cache.

They'll route.

They'll measure.

They'll kill useless agents.

And they will eventually stop asking how many tokens their AI system consumed.

They'll ask a much harder question:

What did we get for them?

Because that is the number that actually matters.

ShareXLinkedInWhatsApp

Ready to Start?

Let's Build Something
Extraordinary.

Whether you have a fully fleshed-out idea or just a spark, at NexKeys we'll help you turn it into a product users love.

48hrResponse Time
100%Transparent Pricing
NDAOn Request