Our season calls come with a scorecard

Before each season we call every player better or worse than his own previous year, then publish how often we were right, by position and age, against what chance alone would score. Further down: the weekly numbers, which are Sleeper's, and the range we put around them, which is ours. Both are graded.

A snapshot measured on 7 September 2026. Updated 28 September; what changed is listed at the end.


Every fantasy tool publishes an order. Almost none publishes a grade.

That is a strange gap, because the grade is cheap to compute and it is the only thing that tells you whether the order was worth reading. A ranking is a claim about the future. A claim about the future can be checked afterwards. Nobody checks.

We started checking, and then we put the results on the page.

What a scorecard actually requires

Grading yourself honestly is harder than it sounds. Three things have to be true, or the number is decoration.

It has to be a call made in advance. Not "our model, re-run today, correlates with what happened." Before each season we mark every player better or worse than his own previous per-game rate. That call is frozen before a snap is played, and it is the thing being graded.

It has to be precision, not recall. The tempting question is "of the players who improved, how many did we catch?" A system that marks every player up scores 100% on it. The honest question is the reverse: of the players we called, how many did what we said? We make the up-call on only 36.2% of players, fewer than actually improve, and that selectivity is what stops the number being optimism.

It has to be measured against chance. Only about 43% of players improve in a given year, so a correct "better" is the harder of the two calls. And the base rate is not constant: young players improve far more often than old ones, so a single league-wide baseline flatters every young group and damns every old one. Every row below is compared against its own group's rate.

The scorecard

3,981 player-seasons across six league chains and four season transitions: players ranked inside 2.5× what a league actually starts at their position.

calledrightchance
"He will be better"1,43969.6%42.6%
"He will be worse"2,53872.7%57.4%

Both calls beat chance, and the harder one, "better", beats it by more. The pooled figure hides the part worth knowing, though: where the calls are strong and where they are not.

By position and age

"He will be better" calls that came true, against how often that group improves anyway
Our calls rightThat group's base rate
0%25%50%75%100%QB 30+97.1%QB 30+: Our calls right 97.1%, That group's base rate 34.8% — 34 callsQB 21–2389.4%QB 21–23: Our calls right 89.4%, That group's base rate 53.2% — 66 callsTE 24–2686.4%TE 24–26: Our calls right 86.4%, That group's base rate 52.3% — 81 callsQB 27–2969.6%QB 27–29: Our calls right 69.6%, That group's base rate 37.2% — 23 callsQB 24–2669.4%QB 24–26: Our calls right 69.4%, That group's base rate 42.1% — 62 callsWR 24–2665.5%WR 24–26: Our calls right 65.5%, That group's base rate 38.9% — 200 callsWR 27–2952.6%WR 27–29: Our calls right 52.6%, That group's base rate 26% — 76 callsWR 21–2375.2%WR 21–23: Our calls right 75.2%, That group's base rate 50.6% — 302 callsRB 24–2663.9%RB 24–26: Our calls right 63.9%, That group's base rate 40.7% — 191 callsRB 21–2369.3%RB 21–23: Our calls right 69.3%, That group's base rate 47.1% — 218 callsTE 27–2965.1%TE 27–29: Our calls right 65.1%, That group's base rate 48.1% — 43 callsWR 30+45.5%WR 30+: Our calls right 45.5%, That group's base rate 32.4% — 11 callsTE 21–2369.4%TE 21–23: Our calls right 69.4%, That group's base rate 61% — 98 callsRB 27–2928.6% ▼RB 27–29: Our calls right 28.6%, That group's base rate 35.8% — 28 callsTE 30+0% ▼TE 30+: Our calls right 0%, That group's base rate 30% — 6 calls

Each row is compared with its own group, because young players improve more often anyway. ▼ marks the two groups where our calls did worse than chance.

groupup-callsrightthat group's ratelift
QB 30+3497.1%34.8%+62.3
QB 21–236689.4%53.2%+36.2
TE 24–268186.4%52.3%+34.1
QB 27–292369.6%37.2%+32.4
QB 24–266269.4%42.1%+27.3
WR 24–2620065.5%38.9%+26.6
WR 27–297652.6%26.0%+26.6
WR 21–2330275.2%50.6%+24.6
RB 24–2619163.9%40.7%+23.2
RB 21–2321869.3%47.1%+22.2
TE 27–294365.1%48.1%+17.0
WR 30+1145.5%32.4%+13.1
TE 21–239869.4%61.0%+8.4
RB 27–292828.6%35.8%−7.2
TE 30+60.0%30.0%−30.0

Thirteen of fifteen groups beat their own base rate. The "worse" calls look similar and slightly stronger: thirteen of fifteen positive, topping out at QB 21–23, where 45 "he will be worse" calls were right 45 times.

The best group of all is quarterbacks over thirty: 34 up-calls, 97.1% right, in a group where only 34.8% improve. "Old players decline" does not survive contact with that row.

An earlier draft of this piece argued the opposite, that our optimism fell off with age. It was wrong because it compared every group to the pooled 42.6% instead of to its own rate. Young players improve more often anyway. Measured against the right baseline, the age gradient mostly disappears.

The two groups we get wrong

Running backs aged 27–29: 28 up-calls, 28.6% right, against a 35.8% base rate. When we say a back in his late twenties will improve, you would do slightly better assuming the opposite. That is a real negative result on a real population.

Tight ends over thirty: six up-calls, none right. Six is not a rate, it is an anecdote, and it is printed here as one. We simply don't know whether we are good at this group, because there aren't enough thirty-year-old tight ends to look at.

Both rows stay on the page. A scorecard that only shows the cells you win is a leaderboard.

What the call is built from

The call comes out of a four-step valuation, and each step is there because the one before it is not enough on its own:

  1. What he did last season, under your league's scoring: six-point passing touchdowns, yardage bonuses, whatever your rulebook says. Not generic PPR.
  2. Pulled back toward the opportunity he earned, so a player who outran his usage is not paid for luck.
  3. Blended with the season projection, at a weight chosen by measuring it.
  4. Replaced by the projection when the two sharply disagree: the case where something structural changed and last season is the wrong guide.

The blend beats both of its own inputs. On one league's board, across four season transitions and 2,934 player-seasons, the average miss runs (lower is better):

predictoraverage miss
the blend1.49
the season projection alone1.75
last season alone1.97

That is 24.6% better than using last season alone, and better than the projection it contains. The figure is per league: it runs from 12.0% to 24.6% across the leagues we have ingested. The direction is the same everywhere; the size is not.

The weekly numbers: Sleeper's projection, our range

The projection is Sleeper's

To be clear about whose number is whose: the preseason call above is ours; the weekly projection is not. Start/Sit uses Sleeper's weekly projection, re-scored under your league's own rules. We tried about twenty ways of improving on it, including recent form, usage trends, matchups and expert ranks, and none beat it reliably. So we use it as it is, and grade it the same way we grade ourselves.

Across 25,954 player-weeks, picking between two startable players at the same position:

How often each one picked the higher scorer of two startable players
Sleeper's projection, your rules62.8%bestSleeper's projection, your rules: 62.8% (best)Season average to date61%Season average to date: 61% ()Last week's points57.8%Last week's points: 57.8% ()A coin flip50%A coin flip: 50% ()

25,954 player-weeks, Big Sacks Dynasty. A coin flip gets half right.

predictorpicked the higher scorer
Sleeper's weekly projection, in your rules62.8%
season average to date61.0%
last week's points57.8%
a coin flip50.0%

It is a real edge, and a small one: 1.8 points of accuracy over a plain season average, worth far more on close calls than obvious ones. That 62.8% is one league (Big Sacks Dynasty, measured 4 September); the other leagues we build land at 62.6%, 64.9% and 65.1%, because the rulebook decides which weeks are close. It also pools every position, and about a third of it is defenders and kickers, where the projection has no edge at all.

The range is ours

Sleeper gives the middle number. What we add is the range around it: how high and how low that player's week is likely to land. Every weekly projection on the dashboard carries two ranges, a middle one meant to catch half of his weeks and a wider one meant to catch four in five. They are lopsided on purpose, because a big week can run much further above a projection than a bad week can fall below it.

We fitted the ranges on 2021–24 and graded them on 2025, a season they had never seen:

rangemeant to catchactually caught
middlehalf of weeks51%
widefour in five81%

Both land on target, across 5,925 player-weeks. That is the real test, because one range that happens to fit could be luck. On a close call, compare the ranges, not just the middle numbers: a narrow range is the steadier start, a wide one the bigger gamble.

What the range shows

The width of the range is the most useful thing a projection knows, and it is not the same for everyone. Two terms: a ceiling week is one where a player scores at least one and a half times his projection; a bust week is one where he scores half of it or less.

running back projectedceiling weekbust week
5–8 points22.2%34.9%
11–14 points21.4%18.2%

Moving a back from a low projection to the middle roughly halves the bust risk and costs almost none of the ceiling, in every season from 2021 to 2025. It is a one-sided trade, and the most actionable thing in this article.

At the top it narrows. A quarterback projected 21 or more has a ceiling week only 5.6% of the time, roughly once a season, against about 18% for one projected 11–14. A high projection is not a bigger number so much as a narrower one: past a point, you are paying for certainty with upside.

Now grade the alternative: the experts

Expert consensus rankings are what most managers actually use, and they can be graded the same way. Take the preseason consensus at each position and ask how many of the players it put in that position's top twelve finished in the top twelve.

About half to two-thirds, in most seasons, ranging from a third to over 80% depending on the position and the year. The rank correlation is respectable, 0.74 to 0.76, and it is the number a ranking would quote about itself. The hit rate answers the question you actually asked: a "top 12" call misses on a third to a half of its names.

There is a second finding. For a position's top twelve, when the expert panel disagrees about a player, the consensus misses him by more, most clearly at tight end and running back. That disagreement is a warning you can see before the season starts. Further down the rankings, it tells you little.

None of this says rankings are useless. A 0.75 correlation is real information. It says "top 12" is a weaker statement than it sounds, and that nobody selling you one has told you its hit rate.

What this does not say

  • It does not say we beat the experts at ranking players. That is a different question, not measured here.
  • It does not say the edges are large. Sleeper's weekly projection beats a plain season average by 1.8 points of accuracy; our season call is a real lift on a genuinely hard question. Neither wins you a league on its own.
  • It does not say we out-forecast Sleeper week to week. We don't; we use their weekly number.
  • It does not reach past our data: six league chains, four season transitions, about 6,250 observations. That is our ingested set, not fantasy football at large.
  • It does not fix the two groups we get wrong. Running backs aged 27–29 are a real weakness; tight ends over thirty are unmeasured.

The point

Any of these numbers could be worse than they are and the article would still be worth writing, because the argument is not that our numbers are good. It is that a number you cannot check is not evidence, and the whole category ships rankings nobody checks.

We publish the grade, the base rate it is measured against, the population it came from, and the two groups where we lose. If a competitor's tool is better than ours, that is a fine outcome. But the only way anyone will know is if they publish a scorecard too.


What changed on 28 September

This piece is a snapshot, and four figures were updated as the data moved:

  • Running backs, bust weeks. The table first compared backs projected 0–5 points, a band inflated by games where a back played only on special teams. It now compares 5–8; the finding holds with a smaller gap.
  • Quarterbacks at the top. This said 2.7%. That figure rested on an error in the projection source's stored history, which scored a projected interception as +2 points instead of −1. It is 5.6%.
  • The expert hit rate. This first graded an overall ranking across positions, which in one-quarterback leagues mostly measured where quarterbacks are priced. It read "one in three to two in five". Graded within each position, the consensus does better.
  • Expert disagreement. This first said disagreement predicts a bigger miss everywhere, "up to 1.91×". Measured across all ranks at once, that was mostly depth. Within each position it holds only near the top.

The title and the weekly section were also reworded to make clear whose numbers are whose.

The dashboard recomputes its own versions of these figures every day, so its live numbers will drift from this snapshot. The call-accuracy tables, the outcome distributions and the expert grades are league-independent; the blend comparison is per league and drawn from Big Sacks Dynasty.