I predicted a 467% price hike for AI models in the following months and years (AI is cheap right now. Are you ready for 467% hike?).
This will not happen current development shows.
The main point still holds and gets proven even more over time – it is important that whatever AI management platform your organisation uses, staying independent from the AI model is paramount.
I can now use real insights from working with our customers and the way they decide what AI model to use for what task. And it’s not only the price. The real decision capability of the model is as important and in some cases also a speed.
So the real aim is to follow AI models on the capability to get the job done within a budget, plus adding the awkward fact into equation that the selected AI model capability can change over time.
That’s why we decided to have an AI model job suitability guide to help our customer select the right AI model quickly without the need for long testing.
What the price data show
Still, briefly, let’s have a look how AI model prices developed over time.
Nine AI model providers output token price

$ per 1M output tokens · logarithmic scale · every line carried through to today, dashed where the price is unchanged since the last event. Dotted red rings mark price hikes.
Sources: vendor pricing pages (Anthropic from published API pricing), cross-checked against public reporting and pricing trackers, August 2026
A pattern seems to be forming with latest and improved premium models getting price increases where older or less capable commodity models prices fall down. Overall the costs are going down thanks to collapsing costs and higher efficiency in using AI model tokens.
7 of 10
Providers have raised prices since late 2024.
– 36%
Median price trend per year.
~ 2 400x
Spread between cheapest and dearest model today.
What companies really want
Nobody buys tokens. Companies buy a job finished to a standard and cost they can justify. The main criteria:
- The job must be done properly. A cheap model that gets a contract wrong is not cheap.
- On budget. Not the lowest price, but a known amount we assign. That’s a routing decision to use strong premium models for the silver work and commodity ones for the rest.
- With oversight. Someone has to have the chance to see what was produced and assess its quality and cost.
The 467% hike hasn’t happened yet, but seven of ten priced vendors raised prices and one nearly quadrupled. The median fell 36% a year, but the same job can cost thousands of times more depending on which model you point it at. To act on that information means moving work between models at the cadence the AI model providers moves, which is monthly.
Augela AI model index – capability with costs to complete the work

Artificial Analysis Intelligence Index v4.1 (0–100) plotted against the dollar cost of running that index · logarithmic cost axis · dot size = output tokens per second · solid dots sit on the efficient frontier, faded dots are beaten by something cheaper.
Source: Augela customer’s experience. Artificial Analysis, Intelligence Index v4.1 (Agents 34%, Coding 24%, Scientific Reasoning 24%, General 18%), captured August 2026. Human-preference rankings referenced below from Arena (formerly LMArena), same capture date.
Which model for which job
The capable model tells you what’s rational in general. Not what is rational for your task, because you weight accuracy, speed and cost differently. This is the current quarterly snapshot:
| Job | What actually dominates | Pick today | Cost | Speed |
|---|---|---|---|---|
| Classification, routing, tagging high volume, narrow task, low accuracy bar | Cost, then speed. Accuracy above ~35 is usually plenty. | Recommended GPT-5.6 Luna · medium | $0.01 | 162 t/s |
| Real-time chat, autocomplete a human is waiting | Speed above all — latency is the product. | Recommended Gemini 3.7 Flash · low | $0.16 | 324 t/s |
| Bulk summarising, extraction large documents, moderate accuracy bar | Capability per cent. Nothing beats this point. | Recommended DeepSeek V4 Flash · max | $0.11 | 99 t/s |
| Customer-facing drafting tone and judgement matter | Human preference, not benchmarks — Arena Elo is the better guide, and Claude leads it. | Arena-led Claude Opus family | $0.72+ | ~50 t/s |
| Agentic coding, multi-step work errors compound across steps | Accuracy dominates — a failed run costs more than the tokens saved. | Recommended Grok 4.6 · high | $0.84 | 58 t/s |
| Hard reasoning, research correctness beats cost | Top of the frontier — but stop at xhigh, not max. | Top tier Claude Opus 5 · xhigh | $1.80 | 51 t/s |
| Current-events answers, cited sources — freshness beats reasoning | Live retrieval. No frontier model can substitute, whatever it scores. | Perplexity Sonar Pro | $15/M out | 123 t/s |
| Regulated or on-prem work — data cannot leave the building | The boundary, not the benchmark. Capability is what you can afford to give up. | Self-hosted open-weight via Ollama | Hardware, not tokens | ~110 t/s |
Cheaper AI models on average and simultaneously more expensive at the top is not a contradiction. You face a dynamic routing problem what job to direct to what AI model to keep up with changes.
Augela helps you control all of this
Augela is an AI control layer between your company and whatever model you use. Those three requirements become things you can check rather than hope for.
Done properly → human-in-the-loop review. Review queues put a person between a model’s output and your customer. That’s what sets the accuracy bar for a given job, and it’s what makes changing the model underneath a decision you can take.
On budget → per-agent ROI tracking, your own providers. Connect OpenAI, Anthropic, Mistral, Perplexity, your own self-hosted endpoint or others under your own terms and job needs. Per-agent cost and return is how you can find the agent burning premium tokens on a job commodity AI model can do.
With oversight → audit trails and a knowledge hub. Your company knowledge, guardrails and accumulated corrections and insights live in Augela. That’s the €€€$$$ saved if you want or need to change AI models for your workflows.
Three layers are worth setting up from the start:
- Roles, so not everyone can change everything. Admins configure the platform. Reviewers assess answer quality. Tutors decide what becomes shared knowledge. Everyone else just uses the AI assistant.
- Guardrails and watch words. Define rules the agent must always follow, and flag responses containing sensitive terms, competitor names, or anything else you want a human to see. Governance can run passively (flag and log) or actively (warn, hold, or block a response before it goes out).
- An audit trail you didn’t have to remember to switch on. Every AI decision and admin action recorded in a tamper-evident, hash-chained log with the reasoning, the confidence, and the cost. Exportable to CSV or JSON with chain verification, and streamable to your SIEM.
Ready to get started?
Create an account or get in touch, start a 7-day free trial , no credit card required.
Access the complete human AI interface with transparent monthly payments or contact us to create the optimal package for your business.
Subscribe to Augela Blog
