skip to content

Try Widely, Build Narrowly


Whether to use language models for every part of life is the wrong question, because it treats trying and building as one act. They are not. A try that stays a try costs minutes and leaves nothing behind; building a tool around a task is what costs, and the bill arrives later as maintenance. So the rule is four words. Try widely, build narrowly.

I got there the slow way, through an answer that was too narrow and three corrections from the model I was arguing with. The narrow answer was a bottleneck test. On that test a model pays off where an existing task is limited by reading, drafting, recall, or judgment that has to be exercised at volume, and where a mistake is cheap to catch. Most of life is not limited that way. Sleep, training, and the people you love are limited by presence and time, and a model pointed at them gives you a tool to maintain rather than an outcome that moves. All true, and only half the picture, because the test can only see tasks that already exist.

The other half is the tasks that did not. A model also changes an outcome when it lowers the cost of something until it becomes worth doing at all. Several things I now record or make as a matter of course were never recorded or made two years ago, not because they were bottlenecked but because capture cost more than the result was worth. No bottleneck test could have found them, because there was nothing to measure. They were found by trying, and this is where Ethan Mollick’s first rule in Co-Intelligence, always invite AI to the table, earns its place. The frontier is jagged. The bottleneck test can tell you where a model will not help with a task you already do; it cannot tell you which tasks that do not yet exist would fall under the line, and for those the only way to learn is to bring the model to things you have no reason to expect it can help with and watch.

Side by side those two halves look like a contradiction, restraint against an open invitation. They are not, once you price the two acts separately. Trying is how the new tasks are discovered, and wherever the result is quick to check it is close to free; its real cost is the temptation it creates, because every try that works asks to become a tool. Building is where the cost lives, and not on the day of building. It lives in the thing having to keep working, stay consistent with its neighbours, and stay in the head of the person who owns it.

I know this because I measured it. In August I ran a census over my own setup, a year of small tools built mostly in the enthusiasm of the moment. Of roughly three hundred, about a third had been called in the preceding month. About a quarter had never been called by anything at all. Every one had looked defensible on the day it was built, the reflex after every incident had been to add one more, and nothing had ever been retired. The trying had been right. The building had been indiscriminate, and no check of each addition on its own merits could have caught it, because each was fine on its own. What catches it is either a count of what is actually being called, run periodically, or a question asked at the point of adding that is not about the addition at all: what does this retire. I had neither, and ran the count a year late.

So I now point a model at anything, including things I am fairly sure it cannot do, because the trying is close to free and the surprises are where the value hides. I build a tool only around a task that has earned it, recurring, high volume, and still cheap to check when it goes wrong, and I let a new tool retire an old one by default, because the scarce resource was never the model’s capacity. It was my attention to what I had built.

One more thing the afternoon taught me sits on the trying side of the line. The highest-value use I have is not scaled at all. It is the strongest model on one consequential question, objecting, supplying the counter-case, naming what the draft missed. Nothing is automated; the value is the quality of one decision. It carries no maintenance load, which makes it the use the restraint least threatens and the one the restraint exists to protect. The rule now lives on my principles page with its failure condition attached: the audit quietly turns into a mandate, and the things that must be kept working start rising while the work they serve does not.

Related by topic
  1. The Harness Is Part of the Token Bill
  2. Summarisation is where the judgment should sit
  3. Masking doesn't declassify