The Compounding of Judgment
We run the world on unrecorded judgment. An agent that writes down every belief, scores it against reality, and keeps the record turns judgment from a perishable service into compounding capital.
Somewhere today, an expert is retiring. A surgeon, a chief underwriter, a grid operator, a loan officer: someone who has spent thirty years deciding what to do in situations that don't come with instructions. There will be a lunch, a speech, maybe a watch. Then they go home, and thirty years of judgment goes with them.
Nobody finds this strange. I find it terrifying.
Their replacement will be smart. She will make good calls and bad ones, and in thirty years she, too, will get a watch. Across those two careers the organization will have made tens of thousands of consequential decisions, and it will have learned almost nothing from them. Not because anyone was careless, but because the reasoning was never written down, never scored, never fed back. The analysis behind each decision evaporated the moment the decision was made. What remained was an outcome in a database row and a feeling in somebody's gut.
This is, when you stop to look at it, the normal state of human civilization. We run the world on unrecorded judgment.
Take stock of what the last few decades have scaled. Compute scaled: my laptop today would embarrass the supercomputers of my childhood. Data scaled: we drown in it. Distribution scaled: a teenager with a phone reaches more people than a 20th-century television network. Even prediction scaled. Language models now produce fluent, often-correct answers to almost any question, at near-zero marginal cost.
Judgment did not scale. The act of deciding well under uncertainty still runs on rare, expensive humans, working one problem at a time, building their analysis by hand, forgetting it, and building it again. Worse: judgment does not compound. A bank that has priced credit for a century is not obviously better at it than it was thirty years ago. A hospital, a ministry, a venture fund, pick your institution, makes the same kinds of decisions a thousand times and banks no transferable improvement, because each act of judgment was a perishable service rendered, not an asset accumulated.
We have built an entire economy on top of the one resource that doesn't accumulate. Imagine if money worked this way: every payment made from memory, no ledger, no interest, every generation starting from zero. You'd call it pre-civilizational. For judgment, we call it Tuesday.
"An intelligence which could comprehend all the forces by which nature is animated and the respective situation of the beings who compose it, an intelligence sufficiently vast to submit these data to analysis, would embrace in the same formula the movements of the greatest bodies of the universe and those of the lightest atom; for it, nothing would be uncertain and the future, as the past, would be present to its eyes."— LAPLACE
Laplace imagined an intellect for which nothing would be uncertain. We will never build that one; the universe has genuine randomness in it, and chaos puts a hard horizon on prediction even where determinism holds. But there is a humbler and, I'd argue, more interesting machine hiding inside that thought experiment, one we can build.
Not an oracle that knows the future. An agent that holds beliefs about the future the way a scientist holds hypotheses: explicitly, quantified, falsifiable. Every estimate a distribution, not a number. Every distribution carrying its provenance: this came from that evidence, weighed this heavily, against that prior. And every belief, when reality finally arrives, scored. No vagueness to hide behind, no rhetorical escape hatch. The Brier score doesn't care how eloquent your reasoning was.
This is a different kind of thing from what the industry is currently excited about. A language model produces opinions: fluent, confident, unaccountable ones, a pundit in a bottle. The agent I'm describing runs an epistemology. It cannot be glib, because glibness is expensive when every claim you make is written down and graded against the future. It's Laplace's demon with the one virtue Laplace left out: honesty about its own ignorance. It tells you not only what it believes, but exactly how much it doesn't know and, critically, what resolving that ignorance would be worth, so it can tell you when the right decision is to go gather more evidence, or to wait.
Foresight, it turns out, is not a personality trait. It's a discipline. And disciplines can be mechanized.
Such an agent learns at three timescales. For each decision, it updates: better priors next time, better-calibrated evidence sources, a sharper sense of which signals were noise. Across decisions, it learns strategy: which modeling approaches work for which kinds of problems, where the tails hide, which assumptions tend to break. And then the outer loop, the one that matters most: it learns how to learn. It authors new modeling tools, backtests them against its entire recorded history, and keeps only what demonstrably improves its score. Its capability surface grows not because someone retrains it, but because it accumulates, like a research group whose members never retire, never forget, and read every paper the group has ever written, including all the failures.
Give it a red-team faculty that attacks its own models before reality bothers to. Give it shadow mode, so it learns even from decisions it didn't make, quietly recording what it would have done and grading itself anyway. Give it the ability to see the portfolio of decisions, not just each case: the ten individually correct calls that are collectively catastrophic.
What you end up with is not a tool that answers questions. It's an institution for making good decisions, compressed into software: the scientific method, double-entry bookkeeping, and actuarial discipline fused into a single agent that holds beliefs with skin in the game.
It's worth sitting with the magnitude of this, because I think we're conditioned by the current discourse to think too small.
When judgment compounds, organizations stop amnesic. Expertise becomes cumulative capital instead of a perishable service, for the first time in history. Every consequential decision leaves the institution smarter, permanently, in a form the next decision can draw on. Small teams inherit the decision quality of institutions a hundred times their size. The bottleneck in medicine, in credit, in energy, in research allocation, in policy, everywhere the scarce input was never information but weighed, honest, accountable judgment, starts to move.
And there is a second-order effect I find even more interesting. Today, most human judgment is unaccountable by construction: nobody can audit a gut feeling. An agent that defends every belief it holds, evidence, weights, uncertainty, score, makes high-stakes decision-making legible for the first time. Regulators can inspect it. Boards can interrogate it. Reality can grade it. In a world about to be flooded with unaccountable machine output, that may end up being the whole ballgame: not the agent that sounds smartest, but the only agent whose honesty is measurable.
Of course, understanding is intrinsically dual use, and a machine that sees further than we do is not automatically a machine that wishes us well. The same ledger that makes it accountable is what makes it governable, assuming we build it in from the start rather than bolt it on after. That part is a choice, and it's ours.
But the direction is set. Somewhere out there, the first judgment that compounds is already being kept. A belief, written down as a number with an error bar, waiting patiently for the future to arrive and grade it.
It will not have to wait long. The future is punctual.