Back to blog
4 min read

Three numbers, not one

methodologyproofstatistics

A success rate on its own is a claim. With its number of observations and its confidence interval, it becomes a measurement — something you can judge, compare and challenge.

That is the difference between "our signals are 59% accurate" and this:

59.0%, 95% confidence interval [55.6 ; 62.4], across 793 resolved decisions.

The first statement tells you nothing. The second tells you how many times we were put to the test, and by how much the true figure might differ from the one we display.

Where that one comes from

From a replay of sixteen years of market history, 2010 to 2026, in which we ran the engine's rules as though they had been live throughout. 90,634 market closes, across twenty-two assets, producing 793 decisions that reached their deadline.

The interval is tight — a little over three points either side. That is what a large number of observations buys you, and it is exactly what is missing from most performance figures you will encounter elsewhere.

That same report also contains what it could not demonstrate, rule by rule: only two formal discoveries — and they stem from a single mechanism — one rule refuted, and sixteen verdicts left undecidable for want of sufficient observations. We publish all four categories, not the first.

What the interval lets you do

It gives you one specific power: knowing when not to believe us.

Take two successive measurements of our engine, before and after an internal fix:

Replay Resolved decisions Accuracy 95% interval
Before 34 67.6% [50.8 ; 80.9]
After 32 62.5% [45.3 ; 77.1]

Five points apart. A press release would write "performance down", or quietly drop the line. The intervals, however, overlap along almost their entire length: with around thirty decisions on each side, that gap is indistinguishable from sampling noise. So we wrote it into the report as such — do not read those five points as a decline.

Both figures stay published side by side. The older one is neither withdrawn nor rewritten: it was true for the system that produced it.

The same reasoning cuts the other way, and that is where it gets uncomfortable for us. Last week one of our internal criteria came back at 100% — across two cases. The matching interval runs from 34% to 100%. In other words, that reading cannot tell complete success from mediocre success. We published it with the caveat spelled out, because a "100%" without its sample size is the most misleading number of all.

The rule that makes this possible

It fits on one line, and it has been in our internal rules from the start: no figure is published without its number of observations and its confidence interval. Publication is unconditional — including when the result contradicts us.

This is not a pose, it is an engineering constraint. A single module computes these intervals for the entire product, and the interface never recomputes one of its own. We use the Wilson interval rather than the normal approximation most spreadsheets offer: the latter produces nonsensical bounds as soon as the sample is small or the rate approaches 0 or 100 — precisely the cases where caution matters most.

Two consequences we accept:

  • A threshold is set before the measurement, never after. We write down and record the criterion, then we measure. Without that, the threshold quietly adjusts itself until the verdict passes, and nobody notices — least of all the person writing it.
  • A refuted rule leaves the product. This month, one of our seasonal rules came back at 29% accuracy across 34 decisions, wrong in all four sub-periods of the sixteen years. We pulled it from the catalogue and published the report that demolishes it.

What you can check yourself

This is the part that matters, and it is what separates a method from a sales pitch.

Every report cited in this article is served online exactly as written, with its method, its numbers and its limitations. The one covering the sixteen-year replay runs to several pages and contains everything that does not work. Our measurement protocol is public, and the verification page gives you the commands to recompute our figures from our raw data and compare them, line by line, against what is displayed.

We are not asking you to trust us. We are giving you the means to check — and when a measurement proves nothing, we are the ones telling you so.

In a field where most of what gets published cannot be verified, that is the only promise we know how to keep.

GeoPulse

Follow the markets with GeoPulse

GeoPulse correlates geopolitical events with financial markets using AI analysis of every event.

Create a free account