← JW Lab About

Abstraction Layer, but Think First

An automated search that records what failed learns nothing that transfers. To compound, it has to record why a kind of thing fails — it has to build abstractions over its own results. Which puts the load on the one component least suited to carry it, because a language model will assert a shared structure between any two findings you put in front of it, fluently, and be right often enough that you stop checking.

The loop, and why it needs to generalise

Begin with the failures, because there are so many of them. A campaign of the loop behind this note runs about a thousand experiments against the stock market and keeps roughly four per cent. Each experiment — a trial — states an economic hypothesis, compiles it into code, replays sixteen years of market history against that code, and scores what comes out. The ninety-six per cent that fail are the interesting inventory, and the question is what a machine can do with them.

A loop that files those failures as attempt 412 did not work is running a very expensive random search. Nothing carries forward. The search space here is combinatorially large — thousands of candidate input fields crossed with every way of combining them — so covering it by enumeration is not a strategy, it is a way of spending a budget.

The only thing that changes that arithmetic is generalisation. If a failure can be filed as measures of this kind fail for this reason, one result closes a family. That is the entire mechanism by which a loop's judgment improves rather than merely accumulates, and it is the difference between a thousand trials and a thousand-and-first informed by them.

The component that has to do it is the wrong one

Here is the problem. Deciding that two findings share a mechanism is irreducibly a judgment about meaning, so it goes to the model — there is no rule to write. And a language model is extraordinarily good at producing that judgment on demand. Show it any two results and it will name the structure they share, in fluent, plausible, well-organised prose, whether or not one exists.

Note what makes this worse than an ordinary reliability problem. If the model were wrong most of the time you would stop trusting it within a week. It is right often enough to be useful, and the invented connections are indistinguishable in tone from the real ones, because they are produced by the same process. Confidence carries no information here.

So the abstraction has to be earned against data before it is allowed to steer. Not evaluated afterwards — a story told after the result is unfalsifiable by construction, since whatever came back is what it explains. Earned in advance, or it gets recorded and given no authority.

That is a design requirement, and it has two halves. You need tests that can actually break a proposed abstraction, and you need something that makes those tests happen. This note is about both. The first half went well. The second is where I cannot make it sound better than it is.

One term, used throughout: a cross-sectional operation is one taken across all instruments at a single point in time — a rank, a median, a quantile computed down the column of everything you could have held that day — as opposed to something computed along one instrument's history. Most of what follows is about claims of that shape.

The habit the whole thing rests on

The project's instructions open with what they call the one habit above all, to be read first, every session. Translated: always dig once more into whether two phenomena that look different are the same underneath, and whether two that look alike are actually different. It claims this is the single mechanism by which the loop's judgment improves — not one clever hypothesis, but the habit of seeing common structure beneath the surface, which makes hundreds of attempts compound instead of accumulate.

Two reflexes follow. The first: when expectation and observation disagree, do not discard either. Rejecting the observation as noise is laziness, and so is rejecting the prior because the literature must be wrong. Assume both are true and find the distinction you did not know existed. The worked example in the instructions is real. A freshness measurement — how recently a thing had happened — amplified the signal in one corporate-action context and killed it in another. Both results were correct. The distinction hiding underneath was whether the quantity measured was a behaviour or a state, which nobody had articulated until the contradiction forced it out.

The second: when several surface-different findings arrive together, suspect a common axis beneath them — and symmetrically, suspect that two findings which look identical are orthogonal. Then the sentence this post hangs on. Whether the distinction or the common axis you have imagined is real is settled with data: do the two sides of the line actually separate, are the residuals actually orthogonal. Stop at a blunt distinction and you begin from zero every time. Find the real structure and one lesson applies in a thousand places.

What the machine actually knows

Against that ambition, the mechanical coordinate system is deliberately, almost comically shallow. A trial's machine-readable position is two things: the set of data fields it touches, and a serialisation of its scoring function's syntax tree. A set of strings and a hash. Themes are resolved by a hardcoded nine-bucket dictionary of well-known factor families — a lookup table, with no model behind it at any point.

This is the load-bearing observation of the note. Every genuinely interesting abstraction in this system lives one level above the machine's coordinates, in model-written prose, with no script behind it. The machine knows two trials touch overlapping fields. It cannot know that two findings are the same bet wearing different clothes. The shallow layer is fast and honest about its scope — but the compounding mechanism sits above it, in a layer whose only quality control is method.

Four diagnostics that discriminate

Each of the four is the same move made in a different dimension: take the thing you believe is responsible, vary it, and state in advance which way the result must move if you were right. Neutralise the dimension you are blaming. Flip the sign. Change the weighting. Ask whether two measurement channels see the same content. The examples below are financial; the shapes are not.

I present them as tests rather than results for the same reason — you cannot check my results, and you can use the tests. What makes each one a test rather than an observation is the sentence written before it runs saying what would count as a failure.

1. Neutralise the blamed dimension and watch the direction of travel

Suppose you blame a result on a tilt along some grouping you did not intend to bet on. The test is to remove that component: subtract each group's mean, keep the within-group variation, re-run. If the effect really was a cross-group tilt, the result should move toward zero. That is the pass, and stating it in advance is what makes the test worth running.

Three other things can happen. The result can get worse — more negative, not less — which falsifies the diagnosis outright: the carrier was within-group all along, and stripping the group means concentrated it. It can collapse to near-zero, which sounds like the pass but is a different finding: the effect was group composition and the within-firm residual is nothing, so there is no clean signal to recover. Or it can survive intact, and that is the trap, because survival invites a victory lap. Surviving does not prove a clean residual: removing cross-group composition leaves any within-group gradient untouched, so if the high end of your measure is systematically the growthier firm inside every group, that bias sails straight through.

0 before the test pass — it really was a cross-group tilt the part you did not mean to bet on is gone collapse — composition and nothing else reads like the pass; no residual is left falsified — the carrier was within-group stripping the group means concentrated it survived — which proves nothing a within-group gradient passes straight through
The axis is the magnitude of the effect, not its sign. Note that two of the four arrows travel the same direction and mean opposite things — one says the unintended component is gone, the other says there was never anything else. How far it moved is the only thing separating them, which is why the naive reading of "it got smaller" is a coin flip.

Four outcomes, one test, and only one is the naive reading. I have hit all four. The fourth is the one that cost me something: a result that survived the neutralisation almost untouched and sat just under the acceptance threshold, which reads exactly like a clean within-group effect that needs one more turn of the handle to clear the bar. It was not. Reading the year-by-year pattern instead of the headline showed the growth-tilt fingerprint underneath, and the only reason it did not get banked is that somebody looked at a dimension the summary statistic had collapsed.

2. Flip the sign

Take a directional signal that lost money and bet the opposite tail instead. Pass: the result flips to a roughly mirror-image gain. The effect was real and you had the sign backwards — embarrassing, easily fixed.

Fail, and the informative kind: the result stays in the same losing band, with a year-by-year curve you can nearly superimpose on the original. Then neither tail was the signal. Both tails hold the same basket — the ranking selects membership in some structural cohort, and any tilt computed on that data selects the same cohort regardless of direction. That is un-fixable by re-signing, re-scaling, or neutralising; the response is to abandon the measure and find a different carrier. I have watched this outcome close an entire family of related quantities in a single trial.

3. Change the weighting scheme

A genuine cross-sectional ranking effect is a statement about order, not about size. Re-weighting the portfolio — from size-weighted toward equal-weighted, or anything between — changes how much the claim is worth, but not whether it is true. So pass is the sign holding; the magnitude can move freely, and usually does.

Fail is a sign flip, and it is a hard falsifier rather than a soft warning. A result positive under size-weighting and negative under equal-weighting was never the effect claimed; it was a handful of very large firms carrying the whole thing while the broad cross-section said the opposite. Large-cap masking, not alpha. There is no repair, only a retraction.

4. Ask whether two channels measure the same content

Many quantities are observable through two channels: a stock and a flow — a balance-sheet level and the income-statement line that feeds it. When one channel is dead it is tempting to try the other and hope for an escape. The rule that emerged is sharper than "sometimes it works". The two channels diverge into distinct effects only when they measure different economic content. When they are one content sampled at two points in a lifecycle they collapse onto one carrier, and the second test re-derives the first result at the cost of an iteration.

Both cases are on record. One pair diverged sharply: one channel captured things acquired from outside the firm, the other things built inside it — different economic objects landing in adjacent ledger lines. Another pair collapsed completely, two views of one compensation phenomenon producing indistinguishable failures with the same year-by-year shape. That closed the axis on both channels at once, but cost a trial a two-minute argument about content would have saved. So the test is a question you answer before spending compute: does the content actually differ?

Read the fingerprint, not the headline

Across all four diagnostics the headline number is the least informative thing on the page. Three unrelated failure causes in this loop landed within a narrow band of each other on the summary statistic and were, at that resolution, indistinguishable. What told them apart was the year-by-year return pattern — specifically the sign in two particular years when market leadership was unusually narrow. Three mechanisms, one headline, three completely different year curves.

That generalises past this domain: a summary statistic makes distinct mechanisms look like one, and what disambiguates them is usually in a dimension you collapsed to produce the summary.

Scope of all of this

Every claim above comes from one domain, one data vendor, one backtest engine, one campaign. The four shapes are stated as general at the head of that section and I still believe that, but the rules they produced are local and I have no evidence about how those transfer. Treat the tests as portable and the conclusions as not.

Falsifiers, written before the test

Several diagnoses here are on record with the prediction committed in advance, which matters because the failure mode of an abstraction is not that it is wrong but that it is unfalsifiable in retrospect — whatever comes back, the story accommodates it. Two examples, in structure only. One said: if this effect were specific to the structural feature claimed, it would post a positive result and its worst years would not coincide with a particular monetary-tightening calendar. It shared that calendar exactly — generic market exposure wearing a structural label, and the pre-written prediction is what made that unarguable. Another said a real operating effect would not crater specifically in refinancing-stress years. It cratered in exactly those years, so the carrier was refinancing exposure and the operating story was decoration.

Usefully, one falsifier failed to fire — the predicted disconfirming pattern did not appear — and that is what let a different, more specific diagnosis survive and go on to steer later trials. A pre-registered prediction that does not fire is evidence too, and a system that records only the kills throws away half the information.

The falsifier that was badly designed

The most honest entry in the record is a falsifier that was simply built wrong. The statistic used to rank names was a mean taken over, at most, two observations per name — dominated by the dispersion of those two draws, so a variance selector rather than a measure of the quantity it was named after. The registered falsifier was therefore testing something other than what it claimed, and the kill it produced was void — not a marginal result, a result about a different question.

Two rules were written down afterwards. De-magnitude the ranker before registering a falsifier on it — if the ordering can be driven by scale or sampling noise rather than content, the falsifier measures the artefact. And pair every tail you intend to read with its own placebo, because a null on one side of a distribution does not bound the other. Voiding your own kill is unpleasant; not noticing you should have is worse.

Prediction prepayment

Out of that came the strongest rule in the system, and it is procedural rather than computational. To claim a theory that explains a result, you must simultaneously register a falsifiable prediction about an experiment you have not run yet — before the result is revealed. An explanation offered without a prediction is still recorded but carries no weight in any later decision: a note, not a belief. Two consecutive missed predictions kill the theory, and the corpse is preserved with the reason, because a refutation is knowledge and deleting it means someone rediscovers the dead end later. The same rule governs changes to the machinery: the prediction is committed before the edit, and a miss reverts it. This is the part I would keep if I had to throw everything else away.

A worked example: a family-level theory — one mechanism claimed across a whole group of related measures — was registered with nine experiments staked on it against a threshold declared in advance. It went zero for nine, and was declared dead with the reason attached. Expensive, and far cheaper than a theory that quietly explains every outcome forever.

The named-abstraction ledger

When several distinct hypotheses start to look like expressions of one mechanism, the model may create a block one level up: a named family. Only under a rule worth quoting in translation, though, because it states this post's thesis better than I have:

A name is not free. It can exist only if it stakes a prediction on an untried target. After-the-fact fitting is always available to a language model, so the value of a name can only be proven by a prediction that hits.

The ledger counts three things, and the definitions are where the care is. A predicted hit requires all of: a trial that declared membership before it ran, that passed the quality gate, and whose return residual actually differs from existing members. That third clause does the work — if the year-by-year fingerprint matches and the holdings overlap, it is the same flesh under a new label and the hit is void. An annexed member is added after the fact, once the result was known; it is recorded and grants exactly zero steering authority. A miss is a declared membership that did not pan out. Two consecutive misses is a death verdict, with a standard phrasing: the essence is narrower than the claim.

That distinction — predicted hit versus annexation — is the whole difference between an abstraction and a story. Both feel identical from the inside.

The loop inventing a connection, and paying for it

Now the counter-evidence. The loop has a forced-novelty mode in which it deliberately skips reading its own belief wiki before proposing. The reasoning is sound: inherited beliefs anchor the model onto ground already mined, and cutting that anchor is how it reaches new territory. The cost is precise, and it has been paid. A field that looks freshly unmined by identifier — but which is the pretax twin of an after-tax line already mapped dead — slips straight through the novelty check, and a full batch of five parallel drafts gets spent re-deriving a known result.

The diagnosis written afterwards is the sentence this note needs: sibling identifiers defeat an unmined-by-identifier heuristic, because the economics, not the identifier, is what has already been mined. The system hard-enforces novelty at the level of the string and nothing at the level of the meaning.

There is more in the same direction. A near-miss where reusing a family name almost overwrote a live block of beliefs. A family that needed four separate rediscoveries before anyone wrote down that it was dead. Two entries in the belief graph explicitly marked as corrections of earlier abstractions — the loop grading its own generalisation and finding it wrong. Those last two are the most encouraging artefacts in the record, and also proof that generalisations were made faster than they were checked.

What enforces any of this

All of the above — the family mechanism, the registered prediction, the falsifier, the ledger and its three counters — lives in text blocks written by the model at commit time. So I went looking for the code that reads them.

Measured

Zero matches. Grepping the entire commit path for the keywords that carry the family mechanism, the registered prediction, the falsifier, and the ledger counters returns nothing. No script parses those blocks. No script increments the counters. No script enforces the two-misses rule. No script blocks a trial that omits a pre-declared falsifier. The aggregation tooling that would do it is listed as deferred future work in the project's own to-do file.

So what enforces "verify the axis before it steers"? Two things. First, the declaration must be committed before execution, and it is non-retroactive: you cannot decide after the fact which family a result belonged to — or rather you can, and the ledger has a name for it and grants it no authority. That constraint is enforced by the ordering of writes rather than by a check. Second, the final verdict is a number behind a firewall the model cannot read and therefore cannot tune toward.

Everything between the abstraction and that number is honour-system prose, executed by the same model that proposed the abstraction.

Why this might be as good as it gets

The temptation is to file that as a defect with an obvious fix: write the script. The fix does not exist. Mechanising the check would require a script that can decide whether two economic stories are the same story — not a missing feature but the hard problem itself, restated. Every substitute is a proxy.

The best proxy is residual correlation between return streams: if two strategies are the same bet, their returns after removing common exposures should move together. That test exists here and is useful. But it sits in the wrong place — a duplicate check after the compute is spent, not a family-membership test before it.

What the proxy does not prove

Low residual correlation between two return streams is evidence that they are not the same bet. It is not proof. Two strategies can express one mechanism through different holdings and decorrelate for reasons that have nothing to do with the mechanism — different rebalancing, different size exposure, different sampling of the same cohort. The converse fails too: high correlation can come from a market exposure neither strategy is about. Sameness of mechanism and sameness of returns are different questions, and the second is only ever a noisy read on the first.

Which leaves the honest position. Pre-declaration plus an unreadable scoreboard is weak enforcement: it stops the two cheapest forms of self-deception — retrofitting the family after the result, and tuning toward a number you can see — and nothing else. It happens to be the strongest enforcement available, and calling it weak is a description rather than a complaint.

In that setting, "think first" is not a slogan hung over the door. It is the mechanism, in the literal sense that there is no other. The abstraction layer is where the compounding happens, and the only thing between a real abstraction and a fluent invention is whether somebody stated in advance what would prove it wrong.

Next: Look-Ahead Bias and Path Dependency — closing a backtest's structural leaks from behind.