What should AI models actually do?
The word of the year is “agentic,” and for good reason. A model no longer has to carry the world in its parameters. It can go look something up, or go do something. And yet the largest models keep getting larger, as if the only way to get smarter is to memorize more.
I keep coming back to the opposite move. Keep the data in external services. Feed knowledge into context when it is needed. Let the model spend its capacity on the part that is actually hard: deciding what to fetch, what to ignore, and how to put the pieces together.
So where is the line between what a model should know and what it should look up?
I spent years on DNS. A resolver that tried to store the internet would be a joke. The whole system works because the thing answering the query is small, fast, and willing to go ask. The intelligence lives in the lookup path, the cache, the policy, and the TTL that tells you when to stop trusting what you remember.
Weights are a cache with no TTL. We bake encyclopedias into them, then act surprised when the encyclopedia is stale, expensive to update, and confident about the wrong year. Retrieval exists. Tools exist. We still train as if the only trustworthy memory is the one we melted into the matrix.
A model should not know your company’s inventory, or yesterday’s incident, or the current price of anything. It should not know the PDF you uploaded last Tuesday because it trained on it in some vague sense. That is a storage problem with a retrieval API. The model should know how to ask, how to notice a conflict, and how to stop when the source is bad.
What does belong in the weights, then?
Judgment. Procedure. Taste about what a good answer looks like. The ability to write a plan, check it, and revise it. A feel for when a tool result is incomplete. Language, math-ish reasoning, the shape of software, the shape of an argument. The things that do not go stale every time someone ships a new version of a product.
There is still a messy middle. Some facts are so common that retrieving them is theater. Some procedures are so specific that no base model should be expected to know them. I would rather be explicit about the messy middle than pretend a 400-billion-parameter blob has solved it.
Take this seriously and a lot of the current race looks like the wrong leaderboard. Lower perplexity on a static crawl is not the same as being useful on Tuesday morning, with a messy inbox, a private wiki, and a production system that can actually break. Useful looks more like:
- A core small enough to run where the data already lives.
- A memory you can inspect, edit, and delete.
- Tools with boring contracts, not magic.
- Evals against the jobs you actually care about, not a general vibe.
I build products where the model is not the product. The product is a six-minute session, or a household widget, or a coach that has to remember what already failed. The moment the model has to be a world oracle, the product gets worse. The moment it gets to be a reliable worker with a filing cabinet, things start to click.
So what should models actually do? Not hold the world. Hold the method for moving through it: know when to look something up, when to act, and when to say “I do not know.” The rest is infrastructure, and we already know how to build infrastructure. We just have to stop asking the model to be the database.
Justin DaCosta builds training systems for Amazon Nova and ships iOS apps on the side.