Isolio

Blog /

The Taxi Meter Is Coming: Why AI Pricing Won't Stay This Cheap

Today's flat, generous AI subscriptions are a customer-acquisition phase, not a stable baseline. Why agentic workloads sit in a different cost bracket entirely, and what architecture protects you when the meter starts running.

Isolio

Strategy

Published 2 September 2026

Right now, a lot of small and mid-sized businesses are paying somewhere between twenty and a couple of hundred pounds a month for AI tools that feel, for practical purposes, unlimited. In a conversation on Isolio's podcast, Tamas Feher and IT consultant Akos Voros argue that this era is a deliberate, temporary phase, and that the businesses building their processes around today's pricing are setting themselves up for a shock.

This is the second article in a series drawn from that conversation. The first argued that every AI project should start from a measurable business benefit rather than a tool choice.

Why it feels free right now

Voros's read on the current moment is straightforward: this is the marketing and customer-acquisition phase of the AI industry, and it's being subsidised by investors pouring extraordinary sums of money into acquiring and retaining users.

He reaches for a comparison that will feel familiar to anyone who remembers the early social media era: since Facebook, we've learned that when something is free, you're the product. AI subscriptions aren't literally free, but he sees the same underlying logic in a twenty-a-month package that, priced against the actual compute it consumes, looks awfully generous.

That generosity, he argues, is not permanent. Somewhere inside that flat monthly fee is a bucket of computing capacity, and that bucket has a real cost behind it. He compares it to getting into a taxi: what you're actually paying for isn't "a ride", it's distance and time, whether or not the fare is currently structured that way.

There are already visible cracks. He points to online communities of heavy AI users comparing notes: someone gets a "pro" tier subscription, and a couple of weeks into the billing cycle the responses noticeably degrade, or the system quietly throttles what it will do, before some users give up and switch to a competing model or tool entirely. His prediction is that the familiar package tiers will start including less, not more, over time, until using AI seriously requires stepping up to enterprise contracts billed by the token, meaning by actual usage, the way a utility bill works rather than a flat subscription.

The reversal of the mobile data story

Both participants frame this as a historically unusual pattern. With mobile data, the industry moved from expensive metered plans, paying per message, per minute, per megabyte, toward flat-rate bundles as the technology matured and competition increased.

AI pricing, Voros argues, is running in the opposite direction. Today's flat, generous packages exist specifically to pull as many users and as many businesses as possible into dependency on these tools as quickly as possible. Once companies have restructured their workflows, and in some cases their headcount, around AI systems, the pricing model shifts toward metered, usage-based billing.

He jokes, only half-jokingly, about companies that laid off ten employees to cut costs and ended up paying a token bill equivalent to the salaries of twenty. It's a punchline right now. He doesn't think it will stay one for long, particularly as the technology becomes normalised over the next couple of years, following a pattern that's already further along among larger companies in the US.

Not all AI usage is the same size

A central piece of the conversation is the difference between using a chat-based interface like ChatGPT or Claude, and running an agentic system built on top of those models. Voros extends his taxi analogy to explain it.

Using a chat interface occasionally, the way a person might ask it to draft a letter or summarise a spreadsheet, is like calling a taxi now and then for a specific trip. An agentic system is a different order of magnitude entirely. It's built to pull in data from potentially thousands of sources simultaneously and process it continuously, not occasionally. It's less like calling a taxi and more like running a transportation network around the clock.

The volume of tasks, steps and processes an agentic system chews through is what separates it structurally from casual chatbot use, and that volume is exactly what drives token costs into a completely different bracket, one that can reach six figures for larger companies operating at scale.

Why good architecture keeps the model small

This is where the conversation gets practical rather than just cautionary. Voros makes the point that in a properly built AI agent, the actual language model does surprisingly little of the work.

Most of an agentic system is still conventional software: established, well-tested coding patterns and business logic that existed long before generative AI, with the AI model plugged in only for small, narrowly scoped, tightly controlled pieces of the process. That's not just about cost, although it is cheaper. It's about reliability. Language models still hallucinate, and a business that lets its operations depend on an unreliable component in an uncontrolled way is exposing itself, and its customers, to real damage. The discipline is to keep the model's job small and well-defined enough that it stays trustworthy.

He illustrates this with a project he worked on at ELTE, building a chatbot to answer student questions. Because the university's system needed the enterprise version rather than a consumer subscription, every single answer had a directly measurable cost. That forced a design discipline: questions with a clear, findable answer, like the location of a specific university building, get answered cheaply and immediately, because that's exactly the kind of question people find irritating to have to search for themselves. But when a query gets complicated, or the internal documentation doesn't have a clean answer, the system is designed to hand off to a human quickly rather than attempt to muddle through, because muddling through is where both cost and hallucination risk spike.

The same logic applies to something like processing an invoice with a hundred line items and a typo in it. Older, rigid systems used to reject the whole document on a string-matching failure, while a well-designed AI-assisted process can recognise the typo semantically. But a business still has to decide deliberately which parts of that process get full automation and which get escalated to a person, because at a few pence per line item, human review can already be cost-competitive with an AI process that's trying to parse something ambiguous.

Insurance against price hikes: don't marry one vendor

The other structural defence Voros describes is multi-vendor management, orchestrated through what the industry calls an agentic platform.

If a business builds its automation directly and rigidly around a single model or provider, it has no leverage when that provider raises prices or throttles quality. If instead a business runs its agents through an orchestration layer, it can compare model options and route to whichever one is currently cheaper for a given task. If a provider then raises prices, the business can simply reroute to a different model underneath, often without the rest of the system, or the end customer, noticing anything changed except a lower bill.

He compares the risk profile of running such a platform to running any standard business application, not something exotic, and points out that the same swappability applies component by component: if the text recognition engine in a pipeline gets pricier or worse, it can be swapped out without touching the rest of the chain.

Tamas adds a related point about the trajectory of raw model costs. Newer, more capable models tend to bring down the cost of solving a well-defined, narrow task even as headline prices for cutting-edge capability rise, because a smaller, cheaper model may now be able to do a job that used to require an expensive one.

The net effect, both agree, is that the true cost of running an AI-based process in two or three years is genuinely hard to predict, and businesses shouldn't build their forecasts on today's promotional pricing holding steady. As Voros puts it, the pricing story is still cheap now. The real question is what it costs once the promotional period ends, and building an architecture that isn't hostage to a single vendor's next price change is the closest thing to insurance a business has against that.

The final article in this series looks at the controls that keep that bill from arriving as a surprise in the first place.


This article is drawn from a conversation on Isolio's podcast between Tamas Feher (Isolio) and Akos Voros, an IT consultant whose career spans IBM and ELTE's Applied AI Center.

Related articles

Ready to embed AI inside your product?

Book a Call