skip to content

Principles

This page holds the few rules I actually work by. The entry bar is the same one the library uses: a rule earns its place when it has changed how I work and there is an artefact behind it, and it stays only while it keeps holding. Each entry names the condition under which it fails, because a principle without its failure condition is a slogan.

Spend judgment once and let cheaper systems reuse it, but only while the owner rules and the rules stay tested

The compounding benefit of working with language models does not come from the strongest model answering every question. It comes from spending judgment once and letting cheaper systems reproduce it: skills, memory, checklists, and lately a smaller model learning to score my reading the way I would. Attention paid this way is paid once and reused many times.

The compounding holds only under two conditions, and both are easy to skip. The signal that calibrates the cheap system has to come from the owner, not from the strong model that supervises it. Two models agreeing looks like calibration and proves nothing; the cheap one has simply learned the strong one’s errors. And the taste has to be stored as tests, examples with the expected outcome, rather than as an ever-longer prompt. A prompt that grows by defensible additions drifts, and no per-request check can see the drift, because each addition looked right on its own.

The current worked example is a small model learning to rank my daily reading, calibrated against my rulings on the items where it and the supervising model disagree, with every ruling kept as a fixture that reruns before each review. It is in progress. This entry stays provisional until the loop either earns its place or falls out.

The failure condition is worth its own sentence. A cheap judge earns its place only on a stream whose items must be read and whose ranking is a matter of taste; a threshold on a number wants deterministic code, and a model reading it adds narrative rather than information. And because the scarce input to any such loop is the owner’s rulings, not the model’s tokens, only a couple of loops can be in calibration at once, with a new one added only when an earlier one has been promoted or shut. The lens fails the moment “intelligence on tap” is read as licence to point a model at everything.

Decide where a model changes the outcome one task at a time, and leave the rest of life uninstrumented

The question I get asked, and kept asking myself, is whether to think about using language models for every part of life. The answer is no, and the reason is not modesty about the models. A model pays off where a task is bottlenecked on reading, drafting, recall, or judgment that has to be exercised at volume, and where a mistake is cheap to catch. Most of life is not bottlenecked that way. Sleep, training, family time, and the relationships that matter are limited by presence and time, and a model pointed at them adds a tool to maintain rather than an outcome that moves.

That test has a second half, and the first version of this entry missed it. A model also changes the outcome when it lowers the cost of something until it becomes worth doing at all. Several things I now record as a matter of course were never recorded two years ago, not because the recording was bottlenecked but because it did not exist: capture cost more than the record was worth, and a model moved it under the line. So the question is not only which existing tasks a model speeds up but which tasks it makes cheap enough to exist. Ethan Mollick’s first rule in Co-Intelligence, always invite AI to the table, is the same point from the discovery side: the frontier is jagged, and the only way to learn which tasks fall under the line is to try. The two halves fit because trying and building are different acts. Trying a model on a task costs almost nothing and is how the new tasks are found; building a tool around one is what carries maintenance load. Try widely, build narrowly. The restraint below is about what gets built, not about what gets tried.

There is a third way a model changes an outcome, and it is neither throughput nor cost. On a single consequential question the strongest model is a thinking partner: it objects, supplies the counter-case, and names what the draft missed. Nothing is being scaled; the value is the quality of one decision. This entry was itself corrected three times in one afternoon that way. It sits almost wholly on the trying side of the line, since a conversation carries no maintenance load, which makes it the use the restraint least threatens and the one the restraint exists to protect.

The cost that argues for restraint is not tokens. It is maintenance. Every tool added to a personal system has to be kept working, kept consistent with its neighbours, and kept in the head of the person who owns it, and that load is invisible at the moment of adding because each addition looks defensible on its own. The entry bar for this page is an artefact, and the artefact here is a census I ran over my own system in August 2026: of roughly three hundred small tools built in a year, about a third had been called in the preceding month, and about a quarter had never been called by anything at all. The reflex after every incident had been to add a component, and nothing had been retired. The useful unit of decision turned out to be a short list of recurring, high-volume tasks, reviewed periodically, with the default that a new tool retires one.

This entry is the other half of the one above. That one limits how many judgment loops can be in calibration at once, because the scarce input is the owner’s rulings. This one limits the domains, because the scarce resource is the owner’s attention to what has been built. Both say the same thing from different sides: intelligence on tap is cheap, and the owner’s attention is not, so the model goes where attention is already being spent badly, or where a worthwhile task was never started because it cost too much, and nowhere else.

The failure condition is the one the census caught. The principle fails when the audit becomes a mandate, when “where does a model change the outcome” quietly turns into “how do I put a model on everything”, and the proof is the curve of maintenance load. If the number of things that have to be kept working is rising while the work they serve is not, the leverage has gone negative, whatever each addition looked like on the day it was made.