A Quick Analysis of 2016 Hitter Projections
I’ve been using projections to create dollar values for my fantasy leagues for more than ten years, and even understanding how to convert projections into dollars is just half the battle. The other half is deciding which projections to use in the first place. Should you use only one set of projections? Or multiple? Should you use the freely available projections here on FanGraphs? Or should you pay for projections from other sources? I’m not going to answer any of those questions definitively, but let’s take a look at a handful of projection sources and compare their projections to 2016 actual results.
A few caveats before I start throwing up tables:
1) I chose to use ottoneu FanGraphs points per plate appearance as my primary review, so there is an ottoneu focused lens here, but the linear weights based scoring on a per plate appearance basis correlates very strongly to wOBA
2) I’m not using every projection set that’s available, so this isn’t meant to be an exhaustive review. I primarily focused on the projections available here at FanGraphs, but also included the PECOTA projections from Baseball Prospectus.
3) This covers only 2016 results, so keep in mind that good projection performance last year is no guarantee of future good performance
4) Only players that accumulated 100+ plate appearances and had projections from all systems are included
With that out of my system, let’s take a look at the FGPts Pts/PA root-mean-square error results (note- the lower the number, the better the projections did compared to actual)
| Ages | Count | Steamer | ZiPS | Pod | FANS | PECOTA |
|---|---|---|---|---|---|---|
| 21-26 | 83 | 0.229 | 0.229 | 0.224 | 0.238 | 0.224 |
| 27-30 | 97 | 0.221 | 0.226 | 0.229 | 0.234 | 0.231 |
| 31-40 | 87 | 0.199 | 0.202 | 0.200 | 0.212 | 0.206 |
| ALL | 267 | 0.217 | 0.219 | 0.218 | 0.229 | 0.221 |
For my sample of 267 hitters Steamer had the best Pts/PA projections, but four out of the five sets were very close, with the crowdsourced Fans projections bringing up the rear. For the age 21 to 26 bucket (representing rookies and potential breakout players), the Pod projections (from our own Mike Podhorzer) tied with PECOTA (from Baseball Prospectus). Steamer edged out the other systems in both the 27 to 30 and 31 to 40 age ranges.
So we’ve looked at which systems performed better on a rate basis, but how did the projections do in projecting playing time? Here are the RMSE results for PA:
| Ages | Count | Steamer | ZiPS | Pod | FANS | PECOTA |
|---|---|---|---|---|---|---|
| 21-26 | 83 | 133.5 | 150.8 | 130.8 | 143.5 | 130.6 |
| 27-30 | 97 | 135.3 | 153.0 | 140.9 | 158.5 | 146.6 |
| 31-40 | 87 | 120.7 | 136.8 | 119.2 | 131.5 | 140.8 |
| ALL | 267 | 130.1 | 147.2 | 131.0 | 145.5 | 139.9 |
Steamer once again leads the pack for all players, but keep in mind that the Steamer playing time as shown on FanGraphs is actually fed from our staff maintained Depth Charts, while ZiPS performs poorly here because their playing time estimates don’t account for 25 man rosters or lineup/bench roles.
Looking at the various age buckets, PECOTA and Pod once again do the best with the youngest players, while Steamer and Pod do best with the veterans. It’s interesting to note that every system did worse projecting the 27 to 30 year old players than either other age range.
Overall it seems clear to me that Steamer did the best projecting 2016 performance on a rate basis and playing time (with the help of FanGraphs Depth Charts), but could we have done even better using an aggregate projection approach? Let’s take a look!
| Ages | Count | Steamer | ZiPS | Pod | FANS | PECOTA | Aggregate | Machine |
|---|---|---|---|---|---|---|---|---|
| 21-26 | 83 | 0.229 | 0.229 | 0.224 | 0.238 | 0.224 | 0.221 | 0.222 |
| 27-30 | 97 | 0.221 | 0.226 | 0.229 | 0.234 | 0.231 | 0.223 | 0.223 |
| 31-40 | 87 | 0.199 | 0.202 | 0.200 | 0.212 | 0.206 | 0.197 | 0.196 |
| ALL | 267 | 0.217 | 0.219 | 0.218 | 0.229 | 0.221 | 0.214 | 0.215 |
Well, would you look at that! The Aggregate column represents a simple average of all five projections, and the Machine column omits the FANS (machine is a bit of a misnomer, Podhorzer is surely not a host).
The Aggregate projections for all players did better than Steamer alone, and did better than PECOTA and Pod for the youngest age bucket. Steamer retains its edge with 27 to 30 year olds, and the Machine projections performed best for veterans.
Now let’s do the same thing, but looking at plate appearances:
| Ages | Count | Steamer | ZiPS | Pod | FANS | PECOTA | Aggregate | Machine |
|---|---|---|---|---|---|---|---|---|
| 21-26 | 83 | 133.5 | 150.8 | 130.8 | 143.5 | 130.6 | 125.6 | 124.2 |
| 27-30 | 97 | 135.3 | 153.0 | 140.9 | 158.5 | 146.6 | 138.6 | 136.1 |
| 31-40 | 87 | 120.7 | 136.8 | 119.2 | 131.5 | 140.8 | 118.3 | 117.5 |
| ALL | 267 | 130.1 | 147.2 | 131.0 | 145.5 | 139.9 | 128.3 | 126.6 |
The Machine projection aggregate pretty much runs away with the playing time analysis. The reason the Machine does better than the Aggregate here is due to the exclusion of the Fans projections, which performed very poorly with respect to playing time.
Conclusion
Last year I used a combination of Steamer/ZiPS/Pod projections when preparing my personal ottoneu dollar values, and incorporated the Fans projections as part of my playing time estimates. Based on these results I may swap out ZiPS for PECOTA, or at the very least add PECOTA to my aggregate mix. In addition, the assumption I had that the fans do a pretty good job of estimating playing time seems to have been wrong, so I’ll be sticking to a Steamer (Depth Charts)/Pod mix this year.
Keep an eye out for a similar analysis of pitching projections in January!
Justin is a life long Cubs fan who has been playing fantasy baseball for 20+ years, and an ottoneu addict since 2012. Follow him on Twitter @justinvibber.
Cool analysis. But could you please clarify what is included in the aggregate and machine totals?
If I’m reading correctly, aggregate is all the projections he tested, while machine omits the fan projections.
That is correct
Nicely done Justin
Very nicely done. I have developed a quibble with the fan projections: the window of time for them to be done gets smaller every year. Thoughtfully filling them out takes time and extending the period during which they can be submitted and counted would surely help increase the sample size and quality of average projection.
Yes that’s an obvious handicap the FANS projections suffer from, most of the submissions were probably done in January last year, while the other projection systems I pulled from late March.
“The Machine projection aggregate pretty much runs away with the playing time analysis.”
I’m pretty sure an observed edge of 3.5 PA [over Steamer] is not significant for a sample of just 267 players.
True, that was probably poor wording on my part. What I really meant was that two out of the three age buckets and the overall indicated the Machine aggregate was best, and the one bucket it wasn’t best it was a close second.
Well damn, awesome to see that my manual projections could run with the best of ’em. Looking forward to the pitchers, as I always feel like mine differ more significantly than on the hitter side. That could put me way ahead or way behind!
Couldn’t you report a coefficient of determination instead of just the raw RMS error? Without any scaling, it is hard to make sense of the number.
This isn’t a regression analysis. There is no r^2 to report.
Coefficient of determination is not just r^2 from a regression. That is a very specific case of the coefficient of determination of a linear regression.
To get the coefficient of determination, you need only take the mean squared residuals (the square of the values reported here here), divide it by the variance, and subtract it from one.
Coefficient of determination is a measure of variance between independent and dependent variables. We don’t have that here.
The coefficient of determination is a measure of the amount of variance removed by a model. You have a predictive model here – in fact, you have three of them. If you can calculate the squared residuals (clearly you can, they are reported) then you can calculate a meaningful coefficient of determination. It is *not* a measure of covariance (the fact that the coefficient of determination for a linear regression is the square of the correlation coefficient is a convenient accident).
The notion that you have to do a regression to get a meaningful coefficient of determination is erroneous, and in fact is the cause of a lot of shoddy analysis (e.g. looking at correlations of predictive models and the values they’re supposed to predict).
Oblarg >>> The Embassy.
I’m going to backtrack and be more precise. The question posed is how much error is there between two data sets. RMSE answers that. R^2 answers a different question.
These aren’t two unrelated datasets. One set of data is supposed to predict the other.
Simply reporting the RMSE does a poor job of indicating how good of a predictor any of these models are. It does tell us how good they are *relative to each other,* but it’s very hard to make any sense of the numbers past that.
The analysis would be greatly improved if the errors were reported as a coefficient of determination (which is scaled to the variance of the data being predicted) instead of RMSE. Moreover, it would be nice to include the results for a naive model of simply predicting everyone to repeat what they did last year (or averaged over the past n years).
With the caveat that I’m very much an armchair statistician, here are the results:
Pts/PA
Steamer: .377
Aggregate: .373
ZiPS: .354
PECOTA: .351
Pod: .346
FANS: .309
PA
Aggregate: .422
Pod: .409
Steamer (DC): .401
FANS: .369
PECOTA: .350
ZiPS: .239
Thanks!
To be pedantic, did you generate these by the above-mentioned procedure of dividing the mean squared error by the variance and subtracting the result from 1, or did you square the correlation coefficient?
Doing the latter is actually not quite valid, since it will artificially inflate the value of the coefficient by “hiding” a linear regression in the analysis (when really we expect the model outputs themselves to predict the actual values, not some linear regression of the model outputs to predict the actual values).
Here are the RMSE for 2015 performance compared to 2016 as a point of reference for how much better the projection systems did than a simple projection of prior year results:
’15 Pts/PA RMSE: .285 (worst as noted above was .229 of FANS)
’15 PA RMSE: 179.1 (worst as noted above was 147.2 of ZiPS)
’15 Pts/PA RSquared: .172 (worst as noted in another comment was .309 of FANS)
’15 PA RSquared: .170 (worst was .239 of ZiPS)
Very interesting, thanks for doing this.
Great stuff here, Evil Justin! Seriously though. I really enjoyed it. I think you should just change your name to The Machine to really take this article to the next level.
Good stuff! Thanks,this is useful analysis.