Projection Accuracy: Early March Plate Appearances
I’ve always wanted to do a projection analysis, especially at different time points. I had it started in 2020 and everything fell apart with the shortened season. I’m starting off simply today by looking at at-bat projections from March 1st with the Wisdom of the Crowds prevailing.
The reason I chose March 1st was that The Great Fantasy Baseball Invitational (TGFBI) started drafted that day. I pulled all 14 projections that morning. I contacted the paid providers and all but one agreed to have their name associated with the results. They are:
Projections
- Steamer (FanGraphs)
- ZIPS
- DepthCharts (FanGraphs)
- The Bat
- The Bat X
- Davenport
- ATC (FanGraphs)
- Pod (Mike Podhorzer)
- Masterball (Todd Zola)
- PECOTA (Baseball Prospectus)
- RotoWire
- CBS
- Razzball (Steamer)
- Paywall #1
To create a list of players to compare for accuracy, I took the TGFBI ADP (players in demand at that time), selected the top 450 players (30-man roster, 15 teams in TGFBI), and used the hitters. Generally, all the players were projected with the following exceptions. Rotowire and Pods didn’t have a Josh Rojas projection while Pods also didn’t have projections for Jazz Chisholm Jr. and Mike Brosseau. Additionally, CBS was missing several projections. I am blamed for part of it because I forgot to pull designated hitters and they didn’t project as many outfielders. Finally, I just removed the projection for Yasiel Puig.
To determine accuracy, I calculated the Root Mean Square Error (RMSE) for four different sets of values. RMSE is “measure of how far from the regression line data points are” and the smaller a value the better.
- All hitters in the subset (no CBS).
- All hitters with the three Pod and Rotowire missed.
- Removed the three along with several who missed quite a bit of time (Lewis, Mondesi, Trout, Ozuna, Hicks, Rendon).
- Every hitter that I pulled from CBS.
Finally, I created two additional projections. One is the average of the above projections. For the second one, I averaged the results from the best two non-aggerate projections, a found its rank (not included in the one with just the three removed). With all the theatrics out of the way, here are the results.
| Projection | RMSE |
|---|---|
| ATC | 150.3 |
| Razzball | 151.5 |
| Top-2 | 151.9 |
| Average | 152.5 |
| Davenport | 153.9 |
| Steamer | 155.3 |
| Paywall #1 | 155.3 |
| ZiPS | 155.7 |
| Mastersball | 157.0 |
| PECOTA | 158.5 |
| Rotowire | 158.8 |
| DepthCharts | 161.0 |
| Pods | 161.2 |
| Bat | 169.1 |
| BatX | 169.1 |
| Projection | RMSE |
|---|---|
| ATC | 149.0 |
| Razzball | 150.1 |
| Average | 151.2 |
| Davenport | 152.2 |
| Paywall #1 | 153.5 |
| Steamer | 153.7 |
| Mastersball | 155.7 |
| ZiPS | 155.9 |
| Pods | 156.3 |
| Rotowire | 156.4 |
| PECOTA | 157.4 |
| DepthCharts | 159.8 |
| BatX | 168.1 |
| Bat | 168.1 |
| Projection | RMSE |
|---|---|
| Top-2 | 138.8 |
| ATC | 138.9 |
| Razzball | 140.7 |
| Average | 141.0 |
| Davenport | 142.1 |
| Steamer | 142.3 |
| Paywall #1 | 144.7 |
| Pods | 146.2 |
| Mastersball | 146.5 |
| Rotowire | 147.2 |
| PECOTA | 147.7 |
| DepthCharts | 148.8 |
| ZiPS | 148.9 |
| Bat | 157.5 |
| BatX | 157.5 |
| Projection | RMSE |
|---|---|
| ATC | 141.0 |
| Top -2 | 142.7 |
| Razzball | 143.3 |
| Average | 143.4 |
| Steamer | 145.3 |
| Davenport | 145.6 |
| ZiPS | 146.4 |
| Paywall #1 | 146.6 |
| PECOTA | 147.8 |
| Rotowire | 147.8 |
| Pods | 148.1 |
| Mastersball | 148.5 |
| DepthCharts | 151.3 |
| CBS` | 160.0 |
| Bat | 160.2 |
| BatX | 160.2 |
There is a lot to digest here.
- ATC projections come out on top but that should not be a surprise since they are an optimized combination of several projections.
- The boys over at Razzball performed the best of the independents with Davenport projections holding their own.
- The Wisdom of the Crowds performed as I expected by crushing all but a few projections.
- For the rest, I am backing off from making too many judgments. It is just at-bats with many more projections to compare. The BAT and CBS obviously have some work to catch up, but the others are all just incrementally worse than the one listed before it.
And that is it for today. I know it’s just a small taste and sorry.
What I really like is some feedback on the process before I move any further along. If anyone thinks some hitters should be added or removed, here are the hitters I analyzed. Thanks for any suggestions.
Jeff, one of the authors of the fantasy baseball guide,The Process, writes for RotoGraphs, The Hardball Times, Rotowire, Baseball America, and BaseballHQ. He has been nominated for two SABR Analytics Research Award for Contemporary Analysis and won it in 2013 in tandem with Bill Petti. He has won four FSWA Awards including on for his Mining the News series. He's won Tout Wars three times, LABR twice, and got his first NFBC Main Event win in 2021. Follow him on Twitter @jeffwzimmerman.
I think Carty has said that the Bat and Batx don’t project playing time, they just use Depth Charts/Steamer from Fangraphs.
Considering Depth Charts’ lack of success at projecting playing time (see comment below), it makes sense that Bat/Batx ended up at the bottom of the list.
I thought the raison d’etre of Depth Charts was to improve on projections from Steamer and ZiPS by averaging the rates and doing a more targeted and supposedly better playing time projection. If they project playing time worse than either DC or ZiPS, what’s the continuing rationale for Depth Charts’ existence? Will Fangraphs use this analysis in a decision to eliminate Depth Charts?
Great analysis, Jeff.
That was my thought, too. I’d always assumed DC would be an improvement over the individual systems. Maybe that was true in the past, and the projections have improved?
Methinks a year over year analysis would be necessary to determine if this is a general concern or a one year blip.
Also, this is a point in time analysis. To really get a good comparison, you’d have to check back throughout the season and compare the systems’ rolling playing time projections as the season goes on.
What is the actual metric that you are computing the RSME of?
At-bats. As stated in the first paragraph. I should add more of them.
Could we run the numbers on a per plate appearance basis?
Transporting me back to stats class 20 years ago. Take the difference of projection system rate vs realized rate, then weight the absolute error by the square root of the number of plate appearances.
Goobledy-gook for some. But the intention is to figure out which system does rate stats best before moving onto how to best project plate appearances.
Why remove the injured guys? I can understand removing Ozuna. But with the injured players, all of them have injury histories/concerns to some extent. Pulling them from the sample is going to suppress the performance of a system that hedged more than others, which isn’t fair.
Generally don’t like removing outliers when evaluating model performance unless they’re outliers for reasons that couldn’t possibly be picked up in the projections, like Ozuna or, for pitchers, Bauer.
Is there an issue of survivorship bias here?
I think projecting 400 AB for everyone would have beat out some of the projections here. That’s essentially what ZiPS is doing (with a little more thought for established players).
Basically, you won’t get penalized for guessing to high on players who did not play in MLB.
Actually, I think the issue with the DepthCharts is that you did not adjust each system to match the league average. The DC playing time is uniformly too optimistic. The 2021 average for your sample is about 400 AB, whereas the DC average for those same players is over 500 AB.
For fantasy, the ordering is more important than the scale. So, for RMSE, you need to adjust the average of each projection to match the actual average.
See Tom Tango’s explanation: http://www.insidethebook.com/ee/index.php/site/comments/forecast_evaluations/
Good points here. Is this the pseudonymous Mays Copeland of LastPlayerPicked fame? If so: thank you for your service.
To add to this, the best stuff on that Tango article is in the comments. He mentions this with regards to specific methods for Fantasy Baseball purposes:
“Another thing, this doesn’t help for Fantasy forecasts, where playing time IS part of the forecast. Therefore, what you actually want to do is:
1. (actual OPS minus league OPS) * actualPA
2. (forecast OPS minus forecast League OPS) * forecastPA
3. abs(1 minus 2)
For guys who are missing, I think it’s more appropriate to do what Marcel does (gives 200 PA standard), and make his forecasted OPS 100 points under the league OPS.
I would also force the actual OPS at 100 points below the league average. Why do this? Because what we care about is the Fantasy dollars or salaries. And, those have floors. So, ideally, what we are trying to do is combine forecasted OPS and forecasted PA into dollar figures.
If it makes it easier for people to follow/understand that we should convert the OPS,PA figures into dollars first (which have an obvious floor of zero), then do that.”
Yes, weighting by PA will be important for the rest of the stats beyond playing time.
I thought the best part of the post was that it was directed at Nate Silver for doing the evaluation exactly like Jeff did. And then Nate jumped in with the correction!
One suggestion I’d make, which would be for next year’s edition, is to pull the data much closer to the start of the season. On March 1st most systems (Except Rudy, who is a fiend, and ATC, which I think heavily relies on Rudy’s playing time) aren’t doing many updates. In late February when these projections were all finalized, we didn’t know who was going to be on rosters, and there was no information from spring training playing time battles to take into account. For those who draft at the end of March, those are really the most important aspects to nail. I happen to think Rudy is the best at that too so it might not change anything up top, but I think it’d give a fairer shake to guys like Podhorzer and Mastersball who only do sporadic updates in February and are good at interpreting data form Spring Training. Conversely, it would knock down ZiPS which does not take any of the new information throughout spring training into account.
I pulled both. That analysis will come later
Not surprised my March 1 AB projections didn’t show very well here. I don’t take AB projections seriously until we’re well into spring training and pretty much stick with my random guesses that early in March.
Josh Rojas and Jazz Chisholm completely missing from my projections perfectly illustrates that point. I don’t even waste my time projecting a player if we’re not even sure he’ll get even 200 PAs in the Majors. So I wait the next couple of weeks to really make my major PA tweaks and the March 1 numbers are not even very accurate as of that day, since I know I’ll be making additional major updates over the next couple of weeks. It’s a lot of work doing this stuff manually!
Pretty cool work. Interested in seeing more! Thanks
I don’t know if Jeff is interested in re-running his numbers or not, but I did a quick test using the pre-season projections that are still up on FG (so, not exactly the same as his March 1st benchmark).
Using all 258 of Jeff’s hitters, I get the following RMSE for AB compared to the projected average:
1. 129 ATC
2. 135 DepthCharts
3. 136 Steamer
4. 141 ZiPS
5. 149 LEAGUE AVG (400 AB for everyone)
That matches my own expectations: ATC as a smart average would still be on top. The human element of DepthCharts perhaps offers an improvement over Steamer’s algorithm. ZiPS, as the least concerned with playing time, lags behind.
Importantly, all the systems do much better than just guessing the average (400 AB) for every player. That is NOT true for Jeff’s methodology, where 400-for-everyone finishes in the middle of the pack.
This is all assuming that we don’t care if a projection is correctly centered on the actual range of AB. If a system is consistently too high or too low, that error will get washed out when you convert to fantasy salaries or rankings.