Projection System Life Hacks
It’s never too early to prepare for next season, just as you can never have too many articles about the mechanics of projection systems. Well, ok, the second part of that statement is a lie, but it has been awhile since we’ve talked about how we use projection systems.
Our 2016 Steamer projections are already live on all relevant player pages. You’ve probably noticed us referencing them while evaluating various catchers, first basemen, and second basemen. We’ll continue to do so as we move into third basemen this week.
When I see criticism of projections, it usually comes as some variation of “it got Players X, Y, and Z completely wrong.” As projection users, we have to understand that Steamer and other systems are estimating a range of outcomes. We then represent that range with a single statistical line.
The range of projections should follow a standard bell curve. In other words, a player should almost certainly perform within three standard deviations of their projected line. A standard deviation is different for every player due to inconsistent sample sizes. We’ve seen more of Joe Mauer than Aaron Altherr.
In general, as sample size increases, the accuracy of a projection should increase too. Paul Goldschmidt is projected to be the fifth best offensive player per 600 plate appearances. We have 2,648 major league plate appearances informing that projection.
Goldschmidt features a few important qualities that lend stability to Steamer’s expectations. Entering his age 28 season, he’s still in his peak years. Over the last three campaigns, he’s posted a .400 or better wOBA with relatively consistent plate discipline and batted ball profiles.
Steamer still projects him to fall to a .390 wOBA. The decline can be summed up in two parts. First, for a player as talented as Goldschmidt, there are more ways for him to decline than improve. Also, injuries can happen to anybody. Sometimes, players play through those injuries and struggle. That player could have a temporary or permanent loss of skill. If the player lands on the disabled list, that also influences their production. Steamer tries to look at the player’s statistics as a whole, but our hypothetical injured Goldschmidt is now just ordinary Paul Schmidt.
Players with changing skill sets represent an opportunity to exploit projection systems and defeat owners who depend upon them. The difficulty is actually identifying changed players as outsiders. That’s especially true over the offseason when there are few games and misinformation abounds.
Kris Bryant is projected as the ninth best offensive player. We have 650 plate appearances informing the projections plus a brief minor league tenure. Steamer says he’ll have a 83/29/86/11/.271 fantasy line. Since we have a small sample size, we can intuit that the range of possible outcomes is much larger than it is for Goldschmidt.
In Summary
There are three points to understand:
- Projections are a single-point representation of a range of outcomes. The likelihood of those outcomes resembles a bell curve.
- A change in talent, temporary or permanent, can create arbitrage opportunities.
- Sample size affects the range of possible outcomes.
If you can internalize these three bullets, you’ll have a solid grasp on the strengths and weaknesses of projection systems. Remember, only history can possess a single truth. The present and the future can only be estimated.
You can follow me on twitter @BaseballATeam
Parsing your words in In Summary, I read “three outcomes truth.” So basically this is an endorsement of Joc Pederson?
This would explain why Papi has been easy to get (relatively) 4 out of the last 5 years, especially this year. Even if you would have asked me “do you think David Ortiz is Opsing around .850 to .900” I’d have expected a regression to around .250 25 with a .750 OPS
I know this is written for fantasy but this is a main-blog worthy message and a concise breakdown of concepts many folks fail to consider. Great stuff.
Some corollaries:
1. Rates = good , Counting Stats = bad
The weakest point of projection systems is PA/innings. Don’t obsess over who is projected for the absolute most HR or SB….infer who will hit them out/swipe them at the highest Rate per PA or GP and then do a little mental math on your level of agreement with each individual playing time projection.
2. Skills not Composites
Who cares who is projected for the best ERA/xFIP/WHIP? The error in those numbers folds in all kinds of randomness from the underlying projected components. Focus on skills in projections like K, BB, contact rates, FB and GB %’s. They’ll inform likely composite stat outcomes better than the numbers that are propagating the errors of all of them by multiplying/dividing them by one another.
3. Old Gold (pet theory of mine)
Players that are really old and still really good are outliers but exist. Projections don’t project outliers and beat everyone with the age stick. Especially HOF level talents don’t age the same as the average player. if anything this player class is basically all or nothing and you rarely get the statistical mean of a vintage campaign and collapse the way the numbers suggest.
Is there a way to include some measure of standard deviation in the projections to get a sense of how far the projections may drift?
-T
Buy a subscription to BP. PECOTA projections show dispersion.
As Blue notes, PECOTA does something like this with percentile rankings. However, there is still one serious issue with the approach. The percentile rankings are based on a single true talent level. The 75th percentile stat line just means the player was more than one standard deviation better than expected. It’s a form of luck. If you think PECOTA has pegged the talent level wrong, then all of the percentiles are wrong. If you think it’s right, then you should reference the 50th percentile ranking when making decisions.
While it helps to know range, you can make an easy mental adjustment with just the 50th percentile projection and sample size. If Goldschmidt is projected for 30 HR, that can be read as about 25-35. If Bryant is projected for 30 HR, that can be read as about 20-40.
A good uncertainty calculation *should* account for both model error and inherent randomness (what you call luck). I have no idea how PECOTA handles its uncertainty estimation, but your comment about true talent level is not true in general for uncertainty estimation.
Granted, most uncertainty estimates are horrible and don’t always account for everything that they should. Uncertainty estimates tend to be bigger than practitioners would like, so they bias them down in various ways.
I could be wrong since I haven’t been following BPro for a couple years, but my understanding was that they were reporting SDs rather than doing something more complicated like an uncertainty estimate.
Two minor statistical points:
1. 84%, not 75%, corresponds to 1 SD above the mean (50%)
2. For a rate statistic and a bell curve ( Gaussian distribution), SD(x/N) = sqrt[(x/N)/N]. Using x = HR and N = PA, for Goldschmidt, SD = sqrt[(30/600)/2648] = .00434 and for Bryant, SD = [(30/600)/650] = .00877. Goldie’s estimate is then 30/600 +/- .00434 or an error of about 8.7% (.00434/.05). Bryant’s is 17.5%. So, we would have a range of about 27-33 HR for Goldie and 24-36 for Bryant. This considers only statistical errors, of course, other usual caveats apply. Brad’s ranges given would probably be more realistic.
If we consider 2SD or a 95% CL, then the ranges are 24-36 and 19-41.
1. Right, but PECOTA references 75th and 90th percentile lines.
2. Thanks for mathing it, I just guesstimated. Good to know that I over-estimated