4 min read

How a hypothesis earns capital

The standard of evidence an idea has to clear before it is allowed to become a position, and why the thresholds are written down first.

Most of the work in a quantitative firm is not finding ideas. Ideas are cheap and arrive constantly: from a paper, from an oddity in a chart, from a conversation. The scarce thing is a process that can tell a real one from a plausible one, reliably, when you badly want the answer to be yes.

This is how that process works here.

The thresholds are written down before the test runs

Before any test is executed, three things are recorded: what result would count as evidence, what result would not, and what would leave the question unresolved.

They are written down first because a threshold chosen after seeing a number is not a threshold. It is a description of the number, dressed as a standard. Nobody does this dishonestly. It happens through a series of individually reasonable adjustments, each of which moves the bar a little closer to where the result already is.

Committing in advance costs nothing when the result is clear, and it is the only thing that helps when the result is marginal. Marginal results are the majority.

Validation happens out of sample

A model fitted and evaluated on the same data will describe that data well. This is not a finding; it is arithmetic. The only informative test is against data the model has never seen, evaluated with the parameters fixed beforehand.

This sounds obvious and is routinely violated in small ways, by choosing a start date that happens to avoid an awkward period, by trying a second specification when the first disappoints, by fixing a bug that only became visible because the result was bad. Each is defensible in isolation. Together they turn an out-of-sample test into an in-sample one with extra steps.

The defense is not virtue. It is keeping the record of what was committed to, and being able to show that the sample used for validation was not touched during development.

Results are intervals, not numbers

An estimate reported as a single number invites a confidence the sample does not support. The same estimate reported as a bootstrapped interval usually reveals that the honest answer is “somewhere in a range that includes outcomes we would act on and outcomes we would not.”

That is a less satisfying sentence, and it is a more accurate one. A wide interval is information: it says the sample is too small, or the effect too variable, to justify a decision yet. Collapsing it to its midpoint discards exactly the part that should have given you pause.

An inconclusive answer is allowed to be the answer

This is the part that is easy to write and hard to hold to.

A research process in which every investigation must produce a verdict will produce verdicts, including where none is warranted. The pressure is not usually explicit. It comes from having spent weeks on something, from the discomfort of an empty result, from the reasonable-sounding thought that surely there is something here.

Treating “inconclusive” as a legitimate, recorded outcome removes that pressure. It also makes the positive results mean more: a process that can return nothing is a process whose “yes” carries information.

Reproducibility makes disagreement productive

Every research run here reproduces: the same inputs and the same code produce identical results, and a run that cannot be reproduced is treated as a defect rather than as noise.

The practical value of this shows up in disagreement. When two runs of the same question differ, there are two possible conversations. One is about whose recollection of the setup is correct, which nobody wins and which is not really about the market. The other is about what changed between the two runs, which is a question with an answer.

The second conversation is only available if the runs are reproducible. Without it, a research disagreement decays into a seniority contest.

What this does not do

None of this makes a strategy work. It is entirely possible to follow every step above and arrive at a well-measured, carefully validated, honestly reported conclusion that the idea is not worth trading. That happens more often than not.

The process is not there to produce edges. It is there to make sure that when something looks like an edge, the resemblance is not an artifact of how it was measured, and to make the cost of finding out low enough that you are willing to keep asking.