How we score ourselves
How we score ourselves
as of · data/predict.py · prediction_calibration
47,272 backtested predictions were scored between Oct 5, 2025 and Sep 27, 2026; each was made using only the data available on its own date.
Right means the call's direction matched what happened: a call that leaned higher was right if the price ended higher, a call that leaned lower was right if it ended lower. Coin flips (within 5 points of 50%) are not counted as calls.
The lazy guesses are the bar a call has to clear. Always guessing up is right exactly as often as coins went up; always guessing down as often as they did not; a coin flip half the time. Each is scored on the very same calls as ours, so the gap between the numbers is our edge, or the lack of it.
Range held is the share of outcomes that landed inside the predicted range. The light band is built to hold 8 in 10 outcomes and the dark band half of them, so honest ranges score close to 80% and 50%; much more means the ranges are too wide, much less means too narrow.
The chance score (a Brier score) is the average squared gap between the chance we gave and what happened (1 if higher, 0 if not); lower is better. We compare it with the same score for the base rate — simply using how often coins went up in general at that time. Beating the base rate is the bar for a real edge.
Backtest rows are recomputed every night; live rows are stored the night they are made and never changed, and are scored once their window closes.
Reliability tables
The reliability tables
as of · prediction_calibration.bins
| We said | Predictions | Average chance we gave | How often it went up |
|---|---|---|---|
| <30% | 170 | 25.6% | 73.5% |
| 30-40% | 1,011 | 36.9% | 58.1% |
| 40-50% | 10,924 | 46.4% | 52.1% |
| 50-60% | 3,789 | 52.3% | 54.5% |
| 60-70% | 447 | 64.6% | 74.3% |
| 70+% | 0 | — | — |
| We said | Predictions | Average chance we gave | How often it went up |
|---|---|---|---|
| <30% | 57 | 28.3% | 26.3% |
| 30-40% | 5,780 | 37.4% | 42.2% |
| 40-50% | 9,080 | 44.3% | 39.9% |
| 50-60% | 1,040 | 53.2% | 38.4% |
| 60-70% | 34 | 60.4% | 5.9% |
| 70+% | 0 | — | — |
| We said | Predictions | Average chance we gave | How often it went up |
|---|---|---|---|
| <30% | 2,003 | 28.3% | 47.0% |
| 30-40% | 9,608 | 34.6% | 35.7% |
| 40-50% | 2,969 | 43.6% | 34.9% |
| 50-60% | 360 | 51.5% | 7.5% |
| 60-70% | 0 | — | — |
| 70+% | 0 | — | — |
Coin by coin
Coin by coin: 30-day calls
as of · prediction_calibration per token
Each coin's own backtested 30-day record (327 coins with 20 or more scored predictions). Small samples swing; read these as context, not a ranking.
Caveats
What these numbers do not say
as of
- A high hit rate can come from the market alone: if almost every call says "up" in a rising market, the hit rate equals how often coins rose. That is why each card puts our hit rate next to always guessing up (and always guessing down when coins mostly fell) and a coin flip on the very same calls, and why the chance score is compared with the base rate.
- The backtest runs on the same price history the method was built on. Live scores — recorded the night each prediction is made, never edited — are the stricter test and are shown separately as they come in.
- Windows overlap: weekly 30-day predictions share most of their days, so the 14,940 scored 30-day predictions carry far less independent evidence than 14,940 coin tosses.
- Past patterns can stop working. A pattern that held for years can fail in a new market.
- Prediction from past patterns, not financial advice. Nothing here tells anyone what to trade.