Pre-registration

DWM-1: a pre-registered proof

What we will test, how we will measure it, and what will count as failure — decided before anything runs.

A deep ultramarine field with orange and gold entering from the left, softly grained.

Here is the bet — written down before we could know the answer, and before we could quietly change the question. The hypothesis, the world it runs in, the number that settles it, and the line past which it has failed. What follows is not a result. It is a commitment, published where you can hold us to it.

01

Why we publish the bet first

Almost every lab says its science is rigorous. Far fewer will show you the wager before the outcome is known. Pre-registration is the discipline the rest of science reached for precisely because it costs something: once the hypothesis, the environment, and the success criterion are on the record, the freedom to reinterpret a disappointing run as a quiet success is gone.

So the order here is fixed, and it cannot be replayed. Hypothesis first. Environment first. The measure of success first — and only then the experiment. A conclusion arranged after the fact can always be made to look inevitable; that is storytelling, not evidence. DWM-1 is the first proof of the Large World Model program, run the other way around.

02

The hypothesis

The claim is narrow on purpose, because a narrow claim is one you can break. In a bounded environment with explicit state, a small model trained to predict how that state changes will out-predict a language model of comparable size at what happens next — measured on transitions neither model has seen, under a fixed budget of compute and data.

If the Large World Model thesis has anything in it, this is where it should first become visible: not in how fluently a model describes the world, but in how well it anticipates the world’s next move.

A bet worth publishing is one that can lose. This one is written where it can.
03

The measurement

Success is defined now, in advance: a statistically significant gain in next-state prediction accuracy on held-out transitions, against a language-model baseline of comparable size trained on the same environment. The environment, the split, the metric, and the budget are fixed by this page and by nothing that comes after it.

Nothing is retuned once the first run begins. One environment, one harness, one count — and the count is reported whichever way it lands. The terms below are the terms; they do not move.

Pre-registration record

Sealed before the first run
Hypothesis
A state-prediction model beats a same-size language model at the next transition.
Environment
A bounded world with explicit, inspectable state.
Split
Held-out transitions neither model has seen.
Metric
Next-state prediction accuracy against the language-model baseline.
Budget
Fixed compute and data, frozen with this page.
Result
Empty by design — filled once, after the first run, whichever way it falls.
The record as registered. The result field is empty on purpose — it is filled once, after the run, either way.
04

What failure would mean

If DWM-1 loses, the loss is published here with the same prominence a win would get: the full results, and the gap between what was predicted on this page and what was actually measured.

A negative result is not an embarrassment to be managed. It is the boundary of the thesis, drawn by evidence instead of preference — which is the entire reason the experiment is worth running at all.

Two outcomes, one published record

If it beats the baseline

The thesis holds — for this world, at this size. Nothing more is claimed.

If it does not

The thesis takes the hit — recorded in full, not reframed as a near-miss.

Same page, same prominence. The result is entered once, whichever way it falls.

Both outcomes carry equal weight here, because in a pre-registration they must.
05

Why this should change how you read us

Most labs ask to be trusted on the strength of their victories. We would rather be read on the strength of the method: a lab that binds itself before the run, counts the outcome either way, and publishes the distance between its prediction and its measurement is a lab whose larger claims you can weigh when they arrive.

The ambition behind this program is large. The discipline around it is meant to be larger — and this page, left exactly as it was registered, is where that discipline starts.

Read next