Edgewise logs a probability before the outcome is known, then scores it once the outcome exists. No hindsight, no quiet deletions, no cherry-picked wins. These numbers update automatically and we do not get to approve them first.
2,045 readings over 8 days · calibration grade too few days to grade · calibration error 1.53% · directional accuracy 51.7%
Scope, stated plainly: this ledger scores the direction engine's component probabilities — the chance the next 1h, 4h and 1d candle closes up — logged every hour. The app does not print these numbers themselves, so this grade is of the engine, not of any figure on its pages. Our level-reach and close-location probabilities are not in it yet. When they are, they appear here — with whatever they show.
Correction, 2 October 2026. Every call was graded on the last hourly bar that opened before its resolve time, so a call could run up to an hour past its horizon. Calls are now graded on the last bar finished by their resolve time, and only where that window covers at least 90% of the horizon. The headline now counts only calls whose price at the call and at the outcome both came from Hyperliquid; the 29,122 earlier calls priced on another venue are set apart, not pooled. At the correction: 161 graded calls, directional accuracy 59.63% under both the old and the fixed window, and too few days to grade calibration.
Most trading products advertise a win rate. We think that's the wrong number, and we'll show you ours anyway. What we optimise for is calibration — when we say 70%, does it happen about 70% of the time? That is measured by expected calibration error (ECE): lower is better.
Our raw directional accuracy is 51.7%, and we publish that on purpose. Predicting the next candle's direction is close to a coin flip, for us and for everyone else — anyone claiming otherwise is selling something. What we believe is more useful is how far price tends to travel, which levels get reached and how often, and saying plainly when there is no edge at all. Those are not yet scored in this ledger, so treat that as our claim rather than as something these numbers prove.
Grades are our own scale, published so you can check us: A ≤ 3% calibration error, B ≤ 5%, C ≤ 8%, D ≤ 12%, otherwise F.
| Timeframe | Independent calls | Calibration error (resampled range) | Directional |
|---|---|---|---|
| Loading… | |||
Thin samples are marked. Several FX/index symbols score poorly on small numbers of calls; they stay on this page rather than being quietly dropped. The bracketed range is what the same symbol's calibration error came out at when we resampled its own days 300 times — the honest width of the number beside it. On a small sample that range covers almost the whole scale, which is the point: resampling BTC's own days at ten independent calls returns our worst grade on half the draws, and BTC is one of our best-calibrated symbols. Where the range is wide, the number is not a finding. It usually sits above the point estimate too, because calibration error is biased upward on small samples.
| Symbol | Independent calls | Calibration error (resampled range) | Directional |
|---|---|---|---|
| Loading… | |||
Each row: when we predicted around X%, it actually happened Y% of the time. A perfectly calibrated model has X and Y equal. Rows with tiny samples are noisy by nature — we show them anyway.
| We said | It happened | Sample |
|---|---|---|
| Loading… | ||
A prediction is written to the ledger at the moment it is made, with a timestamp and a horizon. When that horizon closes, the outcome is compared against what we said. Predictions still inside their horizon are counted as pending and excluded — they cannot be quietly dropped if they go the wrong way. Nothing on this page is entered by hand.