We scored 1,581 crypto predictions from YouTube. The bullish ones were worse than random.
Our feed publishes AI summaries of crypto YouTube videos. We asked a simple question: did any of the token calls in those summaries — bullish or bearish — turn out to be right? Every number below traces to a verbatim source quote and an exchange candle. Nothing is estimated, and the misses stay on the board.
Three results survive every robustness check we threw at them: bullish calls — 84% of everything said — landed at worse moments than random dates on the same assets; the bearish calls that look heroic were one cycle top, called once; and the only stable split in the data is who is speaking — people with positions scored at baseline, people with theses scored reliably below it.
What was measured
| Source | AI summaries of crypto YouTube videos — the same items our feed page renders |
| Summaries pulled | 1,230 (full pagination to exhaustion) |
| Channels | 11 |
| Video dates | 2024-04-13 → 2026-07-31 |
| Candidate text blocks | 4,411 (paragraph-level, from 733 summaries) |
| Extracted directional calls | 2,368 |
| Verified as genuinely forward-looking | 1,581 (67%) |
| Primary sample (one view per video × asset) | 673 across 30 assets |
Prices: Binance 1h klines (USDT), Hyperliquid candles for assets Binance does not list (HYPE, MNT). Spot was cross-checked at run time — BTC read $62,933 on Binance, $62,872 on Coinbase, $62,876 on Kraken.
The market backdrop — this matters for reading everything below
| Asset | Start of window | Peak in window | At measurement | From peak |
|---|---|---|---|---|
| BTC | $66,913 (2024-04-13) | $126,011 (2025-10-06) | $62,963 | -50% |
| ETH | $3,233 | $4,935 (2025-08-24) | $1,863 | -62% |
| SOL | $151.62 | $286.24 (2025-01-19) | $72.90 | -75% |
A raw hit rate is worthless in a market like this — which is why every table below carries baselines.
Extraction quality — the part that decides whether any of this is real
The first extraction pass over-collects. Reading its output showed why: it labelled “Ethereum gained 5%” as a bullish call. That is a fact about the past, not a prediction. So a second, independent pass re-read every extracted call against its source block and classified what kind of statement it actually was:
| Statement type | n | Share |
|---|---|---|
| Forecast — a claim about where price is going | 1,344 | 56.8% |
| Description — reports what already happened | 663 | 28.0% |
| Valuation — “undervalued” / “overvalued” | 184 | 7.8% |
| Position — speaker is long/short/adding | 168 | 7.1% |
| Not a call at all | 9 | 0.4% |
The second pass also caught 277 calls (11.7%) where pass one read the direction backwards, and 39 (1.6%) where the sentiment belonged to a different asset named nearby. Everything below uses only forecast + position + valuation statements with direction and asset confirmed: 1,581 of 2,368 (67%). The 787 descriptions are excluded — scoring them would have inflated every number on this page. Provenance: 76% of quotes match the source text character-for-character; the rest differ only by whitespace normalisation. 9 of 4,411 blocks (0.2%) never returned a parseable classification and are simply absent.
The two baselines
A hit rate means nothing on its own, so every result is scored against two controls:
- •Always-long — what you would have scored calling “bullish” on every one of the same (asset, date) pairs. Beating it means the direction mix added something.
- •Random-timing — for each asset, the unconditional probability that a randomly chosen day inside the feed’s own coverage window produced the called outcome over that horizon. Beating it means the calls landed on better-than-random moments. This is the one that tests for skill.
Sample discipline: a video that says “bullish BTC” eight times is one opinion, so the primary sample collapses to one view per (video, asset) — 1,581 calls become 673 views.
Headline scoreboard — 673 verified views
| Horizon | n | Hit rate | 95% CI | vs always-long | vs random timing | p | Median signed return |
|---|---|---|---|---|---|---|---|
| 7 days | 665 | 46.2% | [42, 50] | +3.2pp | -4.1pp | 0.035 | −0.5% |
| 30 days | 625 | 42.6% | [39, 46] | +2.1pp | -8pp | <0.001 | −2.3% |
| 90 days | 508 | 35.6% | [32, 40] | +9.1pp | -10.1pp | <0.001 | −16.7% |
| 180 days | 400 | 26.5% | [22, 31] | +8.5pp | -22.4pp | <0.001 | −26.5% |
| To measurement day | 673 | 25.0% | [22, 28] | +10.3pp | -24.4pp | <0.001 | −20.9% |
Read it this way: the feed beat buy-and-hold (+10.3pp at the final horizon) purely because a minority of views were bearish in a falling market. And it lost to a coin flip on the same assets (−24.4pp). Both are true at once. Applying a 2% deadband (ignore moves smaller than 2%, the same rule our live track record uses) changes no verdict.
Split by direction — and the correction that matters
| Direction | Horizon | n | Hit rate | Random-timing base | Edge | p |
|---|---|---|---|---|---|---|
| Bullish | 30d | 532 | 40.0% | 49.8% | -9.8pp | <0.001 |
| Bullish | 90d | 440 | 28.2% | 43.4% | -15.2pp | <0.001 |
| Bullish | to today | 564 | 14.9% | 47.8% | -32.9pp | <0.001 |
| Bearish | 30d | 93 | 57.0% | 54.4% | +2.6pp | 0.62 (ns) |
| Bearish | 90d | 68 | 83.8% | 60.8% | +23pp | <0.001 |
| Bearish | to today | 92 | 84.8% | 58.8% | +26pp | <0.001 |
Bullish calls were anti-predictive — not merely wrong because the market fell, but wrong relative to picking random dates on the same asset. 84% of all views were bullish. The bearish rows look like skill. They are not, and this is the single most important correction in the study:
The bearish edge does not survive removing one cycle top
All 97 bearish views, scored at 90 days, broken out by quarter:
| Quarter | n | Hit rate | Base | Edge | p |
|---|---|---|---|---|---|
| 2025 Q1 | 10 | 60% | 61% | -1.5pp | 0.92 |
| 2025 Q4 | 33 | 100% | 63% | +37.3pp | <0.001 |
| 2026 Q1 | 19 | 84% | 59% | +25.7pp | 0.023 |
| Everything else | 16 | 50% | 60% | -9.8pp | 0.43 |
2025 Q4 and 2026 Q1 are adjacent quarters straddling the BTC top ($126,011 on 6 October 2025). Strip those two and the bearish edge is negative and insignificant. The honest statement is not “bearish callers have skill.” It is: the feed’s bearish voices got loud at the top of one cycle, once. That is one observation, not a repeatable effect. Do not build a strategy on it — and if you quote this study, quote this box with it.
Did the alt calls at least beat holding BTC?
| Horizon | n | Beat BTC | 95% CI | Median excess |
|---|---|---|---|---|
| 30d | 319 | 44.2% | [39, 50] | −1.9% |
| 90d | 250 | 33.2% | [28, 39] | −7.4% |
| to today | 349 | 42.1% | [37, 47] | −4.9% |
No. Rotating into the feed’s alt picks underperformed simply holding BTC at every horizon.
Explicit price targets: distance decides everything
289 numeric targets were usable (plausible magnitude, direction-consistent). 76 were reached at some point — 26%. That headline hides everything:
| Target distance from spot at call time | n | Reached | Rate | Median days to hit |
|---|---|---|---|---|
| <5% away | 17 | 17 | 100% | 1 |
| 5–15% | 34 | 22 | 65% | 25 |
| 15–50% | 80 | 33 | 41% | 45 |
| 50–200% | 58 | 4 | 7% | 203 |
| >200% (a 3x or more) | 100 | 0 | 0% | — |
Targets within 5% of spot always print — they are not forecasts, they are the current price with a round number on it. Every target asking for more than a 3x has so far failed, one hundred out of one hundred. Only 121 of 289 targets came with any stated horizon at all, so most can never be marked wrong — several are explicitly multi-decade:
“One bitcoin might be worth $6.5 million in the next 20 years.”
By channel — the scoreboard nobody publishes
| Channel | n | Bullish share | Hit 30d | Edge 30d | Hit today | Edge today | Median return |
|---|---|---|---|---|---|---|---|
| Perico | 20 | 35% | 68% | +17.2pp | 65% | +9.9pp | +23.8% |
| Trader XO | 59 | 78% | 53% | +4.4pp | 25% | -17.6pp | −35.6% |
| When Shift Happens | 81 | 95% | 57% | +7.2pp | 20% | -27.8pp | −8.0% |
| Taiki Maeda | 34 | 68% | 35% | -18.8pp | 50% | -10pp | −3.4% |
| Bankless | 83 | 83% | 42% | -7.8pp | 27% | -22.1pp | −25.1% |
| TheRollupCo | 122 | 94% | 38% | -12.3pp | 28% | -20.2pp | −21.6% |
| Invest Answers | 251 | 88% | 38% | -12.9pp | 18% | -32pp | −27.2% |
| ForwardGuidanceBW | 17 | 71% | 31% | -20.2pp | 18% | -36.7pp | −18.0% |
Primary sample, channels with ≥10 views, measured against each channel’s own random-timing baseline.
Only Perico is positive against baseline at both horizons — and it is the only channel whose calls are majority bearish. At n=20 that is suggestive, not proven. The bullish-share column explains most of the table: the more permanently bullish a channel was, the worse it scored.
The calls that actually worked — and the worst
| Date | Asset | Direction | Signed return | Entry → today | Channel |
|---|---|---|---|---|---|
| 2025-02-18 | HYPE | long | +118% | 24.11 → 52.61 | Taiki Maeda |
| 2025-10-07 | XPL | short | +92% | 0.899 → 0.074 | Taiki Maeda |
| 2025-01-27 | SUI | short | +83% | 3.964 → 0.682 | Trader XO |
| 2025-10-13 | XPL | short | +83% | 0.439 → 0.074 | Taiki Maeda |
| 2025-10-15 | ZEC | long | +74% | 262.24 → 455.80 | TheRollupCo |
| 2025-01-27 | SOL | short | +69% | 234.89 → 72.90 | Trader XO |
| 2025-01-27 | XRP | short | +65% | 3.055 → 1.062 | Trader XO |
| 2025-12-17 | DOT | short | +60% | 1.889 → 0.760 | TheRollupCo |
“Shorted around $0.98 to synthetically pre-sell future farming emissions and get paid 50% funding.”
Trader XO’s late-January 2025 short book (SUI +83%, SOL +69%, XRP +65%) is the best cluster in the dataset — it called the alt top within weeks. Worth stating plainly though: the same 27 January 2025 video was bullish BTC (−38%) and bullish ETH (−41%). The shorts were excellent and the longs were not, on the same day. Right on the alts they disliked, wrong on the majors they liked — that pattern is the whole dataset in miniature.
| Date | Asset | Direction | Signed return | Entry → today | Channel |
|---|---|---|---|---|---|
| 2025-02-27 | HYPE | bearish | -152% | 20.92 → 52.61 | Taiki Maeda |
| 2024-11-25 | TIA | long | -96% | 7.761 → 0.319 | Trader XO |
| 2024-11-06 | OP | long | -95% | 1.601 → 0.085 | Trader XO |
| 2025-10-09 | ENA | long | -85% | 0.542 → 0.080 | When Shift Happens |
| 2025-10-07 | MNT | long | -83% | 2.296 → 0.392 | Taiki Maeda |
Is there tradeable intel in here? Four hypotheses, tested
Each tested against the same random-timing baseline, so “the market fell” is already priced out of every number below.
❌ Feed sentiment as a market indicator — it echoes, it does not lead
| Horizon | Forward ρ | p | Backward ρ | p | Verdict |
|---|---|---|---|---|---|
| 7d | −0.04 | 0.78 | +0.11 | 0.49 | nothing |
| 30d | −0.17 | 0.27 | +0.52 | <0.001 | echo only |
| 60d | +0.07 | 0.70 | +0.39 | 0.014 | echo only |
| 90d | +0.22 | 0.21 | +0.32 | 0.07 | nothing |
Daily bull-share index (30-day rolling, ≥8 views) over the dense-coverage span, correlated with BTC forward and backward returns. A plain time index scores ρ=+0.30 on the same test — “it’s just the trend.”
The index tells you where price has been at p<0.001 and where it is going at p=0.27. Twenty correlation cells were tested; exactly one came back significant in the lead direction, which is what twenty tests at α=0.05 produce by chance. There is no sentiment signal here.
❌ Attention as a discovery radar — first mention is a negative signal
| Horizon | n | Beat BTC | 95% CI | Median excess vs BTC | Median asset return |
|---|---|---|---|---|---|
| 30d | 30 | 27% | [14, 44] | −9.1% | −9.7% |
| 90d | 27 | 30% | [16, 48] | −5.5% | −14.4% |
| to today | 30 | 20% | [10, 37] | −18.9% | −50.2% |
Forgetting direction entirely: when an asset enters the feed’s conversation for the first time, it underperforms BTC from that day, at every horizon. This is the attention-top effect — usable as a risk flag, never as an entry. n=30 assets, one regime; exceptions exist (MORPHO +88.7% vs BTC).
❌ “Surprise bearishness” — the messenger does not matter
A permanently-bullish channel turning bearish scored +26.2pp at 90 days (p=0.003); a balanced channel doing the same scored +20.4pp (p=0.010). Statistically identical — who says it adds nothing over what is said. Both numbers also inherit the one-cycle-top problem above.
✅ Traders vs narrators — the one structural finding that holds
| Horizon | Group | n | Hit rate | Base | Edge | p |
|---|---|---|---|---|---|---|
| 7d | Traders | 113 | 55% | 50% | +4.4pp | 0.35 |
| 7d | Narrators | 552 | 44% | 50% | -6pp | 0.004 |
| 30d | Traders | 109 | 50% | 50% | 0pp | 1.00 |
| 30d | Narrators | 516 | 41% | 51% | -9.8pp | <0.001 |
| 90d | Traders | 106 | 52% | 48% | +3.8pp | 0.44 |
| 90d | Narrators | 402 | 31% | 45% | -13.7pp | <0.001 |
Traders: channels run by people with positions (Trader XO, Perico, Taiki Maeda). Narrators: everyone else — people with theses.
Traders sit at baseline — no edge, but never significantly wrong. Narrators are reliably worse than random at every horizon, p<0.01 throughout. The narrowest cut — traders stating an actual position — is the best group in the dataset (90d: +17.8pp, median +27.8%), but at n=19 it is suggestive only. This is the finding to act on, because it is a stable property of who is speaking, not a bet on a market regime.
Caveats, stated plainly
- “To today” is one endpoint, and a bad one. The measurement date sits near the low of a 50–75% drawdown. The 30d and 90d columns are the regime-robust reads; the to-today column mostly measures how long-biased a speaker was.
- 84% of the sample is bullish. The bearish rows rest on 92 primary views.
- Extraction is model-assisted. Two independent passes, every call carries a verbatim quote and a video URL, 76% of quotes match character-for-character. Auditable, not infallible.
- Three assets are 80% of the sample (BTC 324, ETH 120, SOL 93). Per-asset rows below n≈20 are anecdote.
- Survivorship in the source. The feed only summarises videos it ingested; any ingestion bias is inherited here. Coverage is visibly uneven (3 summaries in July 2025, 146 in March 2026).
What this audit changed in the product
The honest summary is that there is product intel here and very little trading intel. The feed cannot tell you where the market is going. It can tell you, with receipts, who was right — and nobody in this category publishes that. So:
- •Every prediction our own AI publishes is already scored on an append-only public track record — the same scoring rules this audit used. A per-channel scoreboard for the feed is next.
- •The summariser is being fixed to separate “what happened” from “what they expect” — 28% of extracted calls were past price action written in forecast-like language, and the share varies by channel (9% to 39%), which is a summarisation artefact.
- •New summaries will emit structured calls (asset, direction, horizon, target, quote) at generation time, so scoring becomes live and exact instead of a two-pass forensic exercise.
- •Channels will be labelled by measured reliability — not hidden, labelled. The traders-vs-narrators split is stable enough to show.
- •First-mention attention now feeds our intel engine as a de-rating input — an asset hitting peak narrative attention is a risk flag, never a buy trigger.
And three things the data says not to do: no sentiment trading signal (it echoes at p<0.001, leads at p=0.27), no inverting bullish calls into a short strategy (the anti-predictive bullish result and the bearish edge are the same one-cycle fact seen from two sides), and no quoting the bearish edge without the quarter table above.
Method, reproducibility, and the dataset
Pipeline: paginate the summary corpus to exhaustion → survey which of the CoinGecko top-1000 assets appear → deterministic candidate extraction over 4,411 paragraph blocks → LLM pass one extracts directional calls → prices attached from Binance 1h klines and Hyperliquid candles → independent LLM pass two re-reads every call against its source block (forecast vs description, direction check, asset check) → scoreboard. Every row carries the verbatim quote and video URL, so any single number on this page can be checked by hand.