How the lab works
Method
Anyone can find a pattern in market data. The hard part is knowing whether you found it or made it. Everything below is machinery for telling those two apart, and for making it costly to fool ourselves.
Hypotheses are cheap; evidence is expensive
A market hypothesis takes a minute to write and a moment to believe. Evidence that it is true takes a clean sample, a specification fixed in advance, a null model that reflects what you actually want to rule out, and the discipline to report the answer you got rather than the one you went looking for.
So the lab treats hypotheses as abundant and evidence as scarce. Ideas are generated freely and registered in a queue; each one that gets run consumes a budgeted slice of a finite sample. That asymmetry is deliberate. It is what stops the lab from spending its data on the fortieth variation of an idea that already failed.
Preregistration and frozen specifications
Before an experiment runs, its full specification is written down and frozen: the universe of instruments, the formation and outcome windows, the estimator, the controls, the null model, the sample period, and the direction the effect is expected to take. Freezing means a cryptographic hash is taken of that specification, and the hash becomes the experiment’s identity.
The hash is the point. A frozen specification cannot be edited after the result arrives without producing a different hash, so “we always meant to test it that way” stops being available as a move. Where an experiment has a frozen specification, its hash is shown on the experiment’s page.
The registered direction matters too. When a result lands on the opposite side of the direction that was registered, that is reported as such rather than being quietly reinterpreted as support for a two-sided claim.
Point-in-time discipline
The most common way a backtest lies is by using information that was not available at the moment it claims to act. Signals are formed inside a window that closes before the outcome window opens, and every input is checked for whether it was genuinely knowable at that boundary, not merely dated as if it were.
This is stricter than it sounds, and it is expensive. One published record here is an experiment that was retired before a single test was spent, because a data field’s intraday vintage could not be established: a snapshot taken at the close and a snapshot taken at the open produce identical histories after the fact, so no amount of retrospective checking can tell them apart. Using it anyway would have been silent lookahead. That result carries an instrument failure disposition, which exists precisely so that a measurement problem is never filed as a scientific finding.
Sessions with incomplete inputs are dropped rather than filled in. Reported sample counts reflect those drops; nothing is imputed or interpolated.
Permutation inference
Significance is not read off a textbook distribution. It is computed by permutation: the signal is repeatedly reshuffled against the outcome (the standard is 999 draws, and the exact count is recorded on each experiment), and the observed statistic is compared against the distribution of statistics the shuffles produce. With 999 draws the smallest attainable p-value is 0.001, so that is the floor you will see reported, never something smaller.
The permutation is constructed to be a null of incremental information, not of raw correlation. Shuffles are restricted so that the controls and the outcome stay paired and only the signal moves, and so that reshuffling respects the data’s own structure over time. The question the test answers is therefore “does this signal add anything beyond what we already control for”: a much harder question than “is this signal correlated with the outcome”, and the only one worth asking.
Discovery is not confirmation
Not every run is an experiment. Smoke runs exist only to check that the machinery works: that the pipeline can see an effect everyone already knows is there. They are never reported as findings. A registered discovery experiment is different: it is preregistered, it is scored, and it is published whatever it finds.
But a discovery result is still a result measured on data the research process has already had access to. Held apart from all of it is a protected-future partition, data the lab has never analysed and will not touch until a prospective confirmation test is registered against it. It is not a hold-out set that gets peeked at; it is the part of the record the research process is structurally unable to have fitted to.
This is why every experiment page carries the same standing sentence: a discovery result is not evidence of alpha; it is a preregistered test result on the discovery substrate. Some results are queued for confirmation, and those pages say so. Nothing has graduated yet. No claim on this site has passed prospective confirmation on unseen data, and none will be presented as though it has.
We publish nulls
A finding that survives is only meaningful if you can see how many did not. When only successes are published, the published record is a selection, and its apparent hit rate says more about the filter than the market.
So every registered experiment is published: the interesting ones, the flat nulls, the results too unstable to rest anything on, and the ones a measurement problem killed. The graveyard is a first-class part of this site rather than an appendix, and a null there gets the same detail (question, design, statistics, limitations) as any other result.
What we never publish
The lab’s market data is licensed from a commercial options market-data vendor, and the terms of that licence are a hard boundary rather than a preference. What crosses to this site is only ever an Acetate-authored, reviewed, sanitized research document.
Specifically, and by construction, this site never carries:
Raw or row-level market data. No quotes, prices, volumes, or option-chain records. Not as files, not as tables, not embedded in a chart.
Vendor identifiers. The vendor is not named here, and neither are its products or interfaces.
Anything granular enough to reconstruct the source data. No time series, no per-instrument tables, no per-date breakdowns. A published experiment reports named summary statistics and prose; the format it is published in has no way to express a series in the first place, so this is a structural guarantee rather than a promise to be careful.
Publication is manual. Every document is reviewed by a person and approved against the exact content that ships, and any published document can be retracted through the same path.
Follow the lab
New experiments are published in batches. One email each time, results and nulls alike.