A Quick Analysis of 2016 Hitter Projections

I’ve been using projections to create dollar values for my fantasy leagues for more than ten years, and even understanding how to convert projections into dollars is just half the battle. The other half is deciding which projections to use in the first place. Should you use only one set of projections? Or multiple? Should you use the freely available projections here on FanGraphs? Or should you pay for projections from other sources? I’m not going to answer any of those questions definitively, but let’s take a look at a handful of projection sources and compare their projections to 2016 actual results.

A few caveats before I start throwing up tables:

1) I chose to use ottoneu FanGraphs points per plate appearance as my primary review, so there is an ottoneu focused lens here, but the linear weights based scoring on a per plate appearance basis correlates very strongly to wOBA
2) I’m not using every projection set that’s available, so this isn’t meant to be an exhaustive review. I primarily focused on the projections available here at FanGraphs, but also included the PECOTA projections from Baseball Prospectus.
3) This covers only 2016 results, so keep in mind that good projection performance last year is no guarantee of future good performance
4) Only players that accumulated 100+ plate appearances and had projections from all systems are included

With that out of my system, let’s take a look at the FGPts Pts/PA root-mean-square error results (note- the lower the number, the better the projections did compared to actual)

FGPts Pts/PA RMSE ’16 Projections vs Actual
Ages Count Steamer ZiPS Pod FANS PECOTA
21-26 83 0.229 0.229 0.224 0.238 0.224
27-30 97 0.221 0.226 0.229 0.234 0.231
31-40 87 0.199 0.202 0.200 0.212 0.206
ALL 267 0.217 0.219 0.218 0.229 0.221

For my sample of 267 hitters Steamer had the best Pts/PA projections, but four out of the five sets were very close, with the crowdsourced Fans projections bringing up the rear. For the age 21 to 26 bucket (representing rookies and potential breakout players), the Pod projections (from our own Mike Podhorzer) tied with PECOTA (from Baseball Prospectus). Steamer edged out the other systems in both the 27 to 30 and 31 to 40 age ranges.

You Aren't a FanGraphs Member
It looks like you aren't yet a FanGraphs Member (or aren't logged in). We aren't mad, just disappointed.
We get it. You want to read this article. But before we let you get back to it, we'd like to point out a few of the good reasons why you should become a Member.
1. Ad Free viewing! We won't bug you with this ad, or any other.
2. Unlimited articles! Non-Members only get to read 10 free articles a month. Members never get cut off.
3. Dark mode and Classic mode!
4. Custom player page dashboards! Choose the player cards you want, in the order you want them.
5. One-click data exports! Export our projections and leaderboards for your personal projects.
6. Remove the photos on the home page! (Honestly, this doesn't sound so great to us, but some people wanted it, and we like to give our Members what they want.)
7. Even more Steamer projections! We have handedness, percentile, and context neutral projections available for Members only.
8. Get FanGraphs Walk-Off, a customized year end review! Find out exactly how you used FanGraphs this year, and how that compares to other Members. Don't be a victim of FOMO.
9. A weekly mailbag column, exclusively for Members.
10. Help support FanGraphs and our entire staff! Our Members provide us with critical resources to improve the site and deliver new features!
We hope you'll consider a Membership today, for yourself or as a gift! And we realize this has been an awfully long sales pitch, so we've also removed all the other ads in this article. We didn't want to overdo it.

So we’ve looked at which systems performed better on a rate basis, but how did the projections do in projecting playing time? Here are the RMSE results for PA:

FGPts PA RMSE ’16 Projections vs Actual
Ages Count Steamer ZiPS Pod FANS PECOTA
21-26 83 133.5 150.8 130.8 143.5 130.6
27-30 97 135.3 153.0 140.9 158.5 146.6
31-40 87 120.7 136.8 119.2 131.5 140.8
ALL 267 130.1 147.2 131.0 145.5 139.9

Steamer once again leads the pack for all players, but keep in mind that the Steamer playing time as shown on FanGraphs is actually fed from our staff maintained Depth Charts, while ZiPS performs poorly here because their playing time estimates don’t account for 25 man rosters or lineup/bench roles.

Looking at the various age buckets, PECOTA and Pod once again do the best with the youngest players, while Steamer and Pod do best with the veterans. It’s interesting to note that every system did worse projecting the 27 to 30 year old players than either other age range.

Overall it seems clear to me that Steamer did the best projecting 2016 performance on a rate basis and playing time (with the help of FanGraphs Depth Charts), but could we have done even better using an aggregate projection approach? Let’s take a look!

FGPts Pts/PA RMSE ’16 Projections vs Actual
Ages Count Steamer ZiPS Pod FANS PECOTA Aggregate Machine
21-26 83 0.229 0.229 0.224 0.238 0.224 0.221 0.222
27-30 97 0.221 0.226 0.229 0.234 0.231 0.223 0.223
31-40 87 0.199 0.202 0.200 0.212 0.206 0.197 0.196
ALL 267 0.217 0.219 0.218 0.229 0.221 0.214 0.215

Well, would you look at that! The Aggregate column represents a simple average of all five projections, and the Machine column omits the FANS (machine is a bit of a misnomer, Podhorzer is surely not a host).

The Aggregate projections for all players did better than Steamer alone, and did better than PECOTA and Pod for the youngest age bucket. Steamer retains its edge with 27 to 30 year olds, and the Machine projections performed best for veterans.

Now let’s do the same thing, but looking at plate appearances:

FGPts PA RMSE ’16 Projections vs Actual
Ages Count Steamer ZiPS Pod FANS PECOTA Aggregate Machine
21-26 83 133.5 150.8 130.8 143.5 130.6 125.6 124.2
27-30 97 135.3 153.0 140.9 158.5 146.6 138.6 136.1
31-40 87 120.7 136.8 119.2 131.5 140.8 118.3 117.5
ALL 267 130.1 147.2 131.0 145.5 139.9 128.3 126.6

The Machine projection aggregate pretty much runs away with the playing time analysis. The reason the Machine does better than the Aggregate here is due to the exclusion of the Fans projections, which performed very poorly with respect to playing time.

Conclusion

Last year I used a combination of Steamer/ZiPS/Pod projections when preparing my personal ottoneu dollar values, and incorporated the Fans projections as part of my playing time estimates. Based on these results I may swap out ZiPS for PECOTA, or at the very least add PECOTA to my aggregate mix. In addition, the assumption I had that the fans do a pretty good job of estimating playing time seems to have been wrong, so I’ll be sticking to a Steamer (Depth Charts)/Pod mix this year.

Keep an eye out for a similar analysis of pitching projections in January!





Justin is a life long Cubs fan who has been playing fantasy baseball for 20+ years, and an ottoneu addict since 2012. Follow him on Twitter @justinvibber.

23 Comments
Oldest
Newest Most Voted
White Jar
9 years ago

Cool analysis. But could you please clarify what is included in the aggregate and machine totals?

TheEmbassy
9 years ago
Reply to  White Jar

If I’m reading correctly, aggregate is all the projections he tested, while machine omits the fan projections.

Trey Baughn
9 years ago

Nicely done Justin

scotman144Member since 2016
9 years ago

Very nicely done. I have developed a quibble with the fan projections: the window of time for them to be done gets smaller every year. Thoughtfully filling them out takes time and extending the period during which they can be submitted and counted would surely help increase the sample size and quality of average projection.

evo34Member since 2023
9 years ago

“The Machine projection aggregate pretty much runs away with the playing time analysis.”

I’m pretty sure an observed edge of 3.5 PA [over Steamer] is not significant for a sample of just 267 players.

Mike PodhorzerFanGraphs Staff
9 years ago

Well damn, awesome to see that my manual projections could run with the best of ’em. Looking forward to the pitchers, as I always feel like mine differ more significantly than on the hitter side. That could put me way ahead or way behind!

Oblarg
9 years ago

Couldn’t you report a coefficient of determination instead of just the raw RMS error? Without any scaling, it is hard to make sense of the number.

TheEmbassy
9 years ago
Reply to  Oblarg

This isn’t a regression analysis. There is no r^2 to report.

Oblarg
9 years ago
Reply to  TheEmbassy

Coefficient of determination is not just r^2 from a regression. That is a very specific case of the coefficient of determination of a linear regression.

To get the coefficient of determination, you need only take the mean squared residuals (the square of the values reported here here), divide it by the variance, and subtract it from one.

TheEmbassy
9 years ago
Reply to  Oblarg

Coefficient of determination is a measure of variance between independent and dependent variables. We don’t have that here.

Oblarg
9 years ago
Reply to  TheEmbassy

The coefficient of determination is a measure of the amount of variance removed by a model. You have a predictive model here – in fact, you have three of them. If you can calculate the squared residuals (clearly you can, they are reported) then you can calculate a meaningful coefficient of determination. It is *not* a measure of covariance (the fact that the coefficient of determination for a linear regression is the square of the correlation coefficient is a convenient accident).

The notion that you have to do a regression to get a meaningful coefficient of determination is erroneous, and in fact is the cause of a lot of shoddy analysis (e.g. looking at correlations of predictive models and the values they’re supposed to predict).

southie
9 years ago
Reply to  Oblarg

Oblarg >>> The Embassy.

TheEmbassy
9 years ago
Reply to  Oblarg

I’m going to backtrack and be more precise. The question posed is how much error is there between two data sets. RMSE answers that. R^2 answers a different question.

Oblarg
9 years ago
Reply to  TheEmbassy

These aren’t two unrelated datasets. One set of data is supposed to predict the other.

Simply reporting the RMSE does a poor job of indicating how good of a predictor any of these models are. It does tell us how good they are *relative to each other,* but it’s very hard to make any sense of the numbers past that.

The analysis would be greatly improved if the errors were reported as a coefficient of determination (which is scaled to the variance of the data being predicted) instead of RMSE. Moreover, it would be nice to include the results for a naive model of simply predicting everyone to repeat what they did last year (or averaged over the past n years).

Oblarg
9 years ago
Reply to  Justin Vibber

Thanks!

To be pedantic, did you generate these by the above-mentioned procedure of dividing the mean squared error by the variance and subtracting the result from 1, or did you square the correlation coefficient?

Doing the latter is actually not quite valid, since it will artificially inflate the value of the coefficient by “hiding” a linear regression in the analysis (when really we expect the model outputs themselves to predict the actual values, not some linear regression of the model outputs to predict the actual values).

Oblarg
9 years ago
Reply to  Justin Vibber

Very interesting, thanks for doing this.

jtevansMember since 2016
9 years ago

Great stuff here, Evil Justin! Seriously though. I really enjoyed it. I think you should just change your name to The Machine to really take this article to the next level.

Scott GilroyMember since 2016
9 years ago

Good stuff! Thanks,this is useful analysis.