Not Every Problem Needs the Best Model
The AI industry is organized around a race to build the most capable model. Most users, however, do not need the most capable model for most tasks. As the cost of computation falls and smaller systems improve, the central question will shift from which model is best to which model is sufficient. The widespread impact of AI may depend less on universal access to frontier intelligence than on making ordinary intelligence cheap enough to use everywhere.
The wrong question
We have spent the first years of the generative AI era asking which model is best.
The question is understandable. Model releases are presented as competitions. Benchmarks create rankings. A new system is celebrated for reasoning more deeply, writing better software, or solving problems that defeated the previous generation. The frontier is where the excitement, investment, and attention gather.
But most tasks are not frontier tasks. Correcting grammar in an email does not require the strongest reasoning system available. Neither does categorizing a support request, extracting a date from an invoice, summarizing a routine meeting, or rewriting a paragraph for clarity.
Using the largest model for every request is like hiring the most accomplished expert in a field to complete its simplest administrative work. The result may be good, but the mismatch is expensive.
The more useful question is not “Which model is best?” It is “What is the least expensive model that can perform this task reliably?”
Intelligence by the token
AI is increasingly bought and measured in tokens. A token is not exactly a word. It is a unit into which text is divided so a model can process it. A short word may be one token. A longer or unfamiliar word may become several. Punctuation and spaces can matter too. [1]
When a model receives a request, it processes input tokens and produces output tokens. The price of using the model is therefore connected to how much information passes through it, how much computation the model performs, and how much output it creates.
This sounds like a technical billing detail. It is actually part of a larger change in how intelligence is distributed. Expertise was once purchased mainly through salaries, consulting fees, institutional access, or years of education. Machine intelligence can be consumed in tiny increments. A person can buy the equivalent of a few seconds of analysis without hiring anyone or owning the infrastructure that produced it.
Tokenization makes intelligence divisible. Divisibility makes it easier to price, route, combine, and embed inside other products. The email button that improves one sentence can purchase only the amount of intelligence needed for that sentence.
Cheaper changes the market
The cost of useful AI computation has been falling quickly. Better hardware matters, but it is only part of the explanation. Models are becoming more efficient. Companies are learning how to serve them at scale. Techniques such as quantization, distillation, caching, and specialized architectures reduce the amount of computation required for a useful result.
The Stanford AI Index has documented steep declines in the cost of querying systems that reach a comparable level of performance. The precise numbers will continue to change, but the direction matters more than any single price. Capabilities that were expensive and scarce are becoming inexpensive and widely available. [2]
When something becomes cheaper, we do not simply spend less on the same quantity. We usually discover more uses for it. Cheaper storage did not cause companies to maintain the same amount of data at a lower cost. It encouraged them to keep far more. Cheaper bandwidth did not merely reduce the price of sending text. It made streaming video and permanent connectivity ordinary.
Cheaper intelligence is likely to follow the same pattern. A company may begin by lowering the cost of one existing workflow. It may end by placing some form of reasoning inside every workflow.
Fit matters more than rank
A model can be less capable in general and better suited to a particular job.
A smaller model may respond faster, run on a local device, consume less energy, and keep sensitive information inside an organization. A specialized model may understand the vocabulary and structure of one domain better than a larger general system. A deterministic piece of software may remain more reliable than any language model when the task follows clear rules.
This is why the email example matters. If the goal is to correct tone or grammar, the extra reasoning ability of a frontier model may add little. The user needs a dependable edit, not a demonstration of the model's full intellectual range.
The same principle applies beyond writing. A hospital may use one system to transcribe a conversation, another to structure the note, and a more capable model only when the case requires complex synthesis. A software product may handle common questions locally and reserve a frontier model for ambiguous requests. The best system is not necessarily one model. It is a collection of capabilities matched to different levels of difficulty.
Routing intelligence
The practical answer is routing. A routing system estimates what a request requires and sends it to an appropriate model.
Simple requests can go to small, fast, inexpensive models. More difficult requests can move to stronger systems. If confidence is low or the consequences are significant, the request can be escalated again or sent to a person. Research on model cascades and routing has shown that carefully combining models can preserve much of the quality of an expensive system while reducing cost. [3] [4]
This sounds similar to how organizations already distribute work. A capable team does not send every decision to its most senior leader. Routine matters are handled close to where they arise. Difficult or consequential questions move upward. The organization works because it reserves scarce attention for the problems that justify it.
AI products will increasingly need the same discipline. The challenge is that difficulty is not always obvious in advance. A short question can require deep expertise. A long document can require only simple extraction. Routing systems will make mistakes, and companies will have to decide whether they prefer the risk of spending too much or the risk of using a model that is not capable enough.
Abundance creates demand
Lower cost will not necessarily reduce the amount of computation AI consumes. It may do the opposite.
When the cost of one AI interaction approaches irrelevance, products can make thousands of calls where they once made one. A system can consider several possible answers, check its own work, monitor a process continuously, personalize an experience for every user, or coordinate many specialized models behind a single request.
This is the paradox of efficiency. Each unit becomes cheaper, but the number of units grows. More efficient engines made travel less expensive and helped increase how much people traveled. More efficient computation made software cheaper to run and helped put software into almost everything. More efficient AI may lower the cost per token while increasing the total number of tokens consumed.
The result will not be one universal intelligence answering occasional questions. It will be layers of intelligence embedded throughout products and institutions, most of it operating without attracting attention.
What cheap does not solve
Lower cost should not be confused with lower consequence.
A cheap model can still make a damaging mistake. A small system used millions of times can create more aggregate harm than an expensive system used rarely. Routing can also hide accountability. When several models contribute to a decision, it may become harder to determine which component failed and who remains responsible.
Cost optimization can create pressure to use a model that is almost good enough. That may be acceptable when rewriting an informal message. It is much harder to justify when the output affects medical care, employment, credit, safety, or legal rights.
The standard should therefore depend on both difficulty and consequence. A simple but high-stakes task may require more verification than a complex but harmless one. The right model is not only the cheapest model capable of producing an answer. It is the cheapest system capable of producing, checking, and explaining an answer to the standard the situation deserves.
The next competitive advantage
Frontier models will remain important. Someone must push the boundary of what machines can do, and some problems will justify using every available capability.
But the next competitive advantage may not belong only to whoever builds the smartest model. It may belong to whoever learns how to make intelligence economical, dependable, and almost invisible.
That means understanding tasks well enough to know which ones are easy, which are difficult, and which are too consequential to automate casually. It means designing systems that can move between local models, specialized models, frontier models, conventional software, and human judgment. It means measuring the quality of the whole process rather than admiring the capability of its most impressive component.
The email assistant is not merely an example of limited ambition. It is also a clue to how AI will spread. The task is common, bounded, and forgiving. It does not need the best model. As smaller systems become capable enough to handle work like this, intelligence becomes cheap enough to appear everywhere.
The frontier will continue to attract attention. Abundance will create the larger change.
Sources & note
Model prices, performance, and architectures change rapidly. The argument concerns the longer-term direction of cost, specialization, and routing rather than the current price of any particular service.
- OpenAI Cookbook. How to count tokens with tiktoken.
- Stanford Institute for Human-Centered Artificial Intelligence (2025). AI Index Report.
- Chen, Zaharia, and Zou (2023). FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance.
- Ong et al. (2024). RouteLLM: Learning to Route LLMs with Preference Data.