Some Words Regarding Baseball Projections
In yesterday’s Corey Kluber article, a commenter pointed out a Steamer quirk – only three starting pitchers are projected for a sub-3.00 ERA. Last season, 22 qualified pitchers finished the year under the 3.00 benchmark. If you set the threshold at 50 innings pitched, 39 starters were below a 3.00 ERA. Clearly Steamer is crazy. Right? I mean, we have to expect a lot more than three pitchers to demonstrate excellence.
Or maybe it’s not so crazy. Steamer provides a single projection based on a range of possible outcomes. Is it hard to believe that most pitchers aren’t likely to post a sub-3.00 ERA?
Let’s get this out of the way. I’m not a Steamer expert, I don’t know the inner workings of this particular system. My goal today isn’t to specifically address Steamer’s quirks, but to speak at a high level about all projections. Remember, you should trust the projections. But it’s useful to understand how they work in general.
You can think of a projection as a representative outcome for a player. There is some distribution of possible outcomes ranging from nothing to perfection. Perfect never happens, while nothing(or worse) happens all too frequently. Below is a hastily constructed illustration. Think of the X-Axis as WAR and the Y-Axis as frequency. Let’s say the graph is titled – 10,000 simulations of Player N for 2015.
This isn’t a perfect graph by any means, but it captures some of the basics. Who remembers the lesson on right-skewed distributions? The mode is the peak, the mean is somewhere to the right of that, and the median is even further right. Which is most useful to us: the most likely single outcome (mode), the average of all outcomes (mean), or the middle-most outcome (median)? You probably want some combination of the mean and mode. I assume each projection system tackles this question in a slightly different way.
So now we have a rough image of possible outcomes, and we know Steamer projects only three pitchers to beat a 3.00 ERA. Let’s turn to a simplistic example. Let’s say we have 30 pitchers who have an equal chance to finish anywhere between a 2.90 ERA and a 3.20 ERA. With that flat distribution of outcomes, we will project all of them to pitch to a 3.05 ERA. We still expect 10 of them to finish below the 3.00 threshold, but we have no means of projecting which ones. Reality is much more complicated.
Steamer probably expects more than three pitchers to finish with a sub-3.00 ERA. It just doesn’t know which ones. If you think about the risk factors involved with pitchers, it’s not surprising that few are projected to be excellent. They all have about a 30 percent chance to land on the disabled list. If they happen to play through injury, their performance can be a lot worse (see A.J. Burnett). They can find ill-luck, bad defense, or ballpark effects working against them. In many cases, one bad day can all but ensure an ERA north of 3.00.
Remember, the projections work. You should trust them. But it’s useful to understand how they work. Every year, some players will be better than expected and vice versa. The job of the projection system is to set an expectation, not to bet which players will be better or worse than that expectation. Actually, that’s your job as the fantasy owner. It’s up to you to decide which pitchers will outperform expectations. As we discussed yesterday, I think Kluber will be one of them.
You can follow me on twitter @BaseballATeam

You should use your knowledge to supplement projections. Start with the projections, if you think they are off, look for tangible changes. If you find something, you can adjust. But if you don’t, stick with the projections.
Exactly. A projection serves as a default where other information is lacking.
But, if you start doing that, you need to adjust ALL players. They all exist within a steamer universe where only a few pitchers have below a 3 ERA, so if you alter one of them to be better it is more valuable than in the real world where lots of pitchers have ERAs below 3. It gets complicated quickly.
I don’t interpret the post as suggesting modifying the projection per se.
Exactly. The projections are fine as a baseline, but if your intention is target specific players blithely trusting them is foolish.
This is one reason why I loved the Marcel projections because it was explicit in how this worked. Projections are not about hitting the nail on the head, they are about grouping similar players into expected tiers. It’s then up to the you, the user, to decide who is likely to perform better or worse than others around him.
A great, and necessary, article!
Nice article. I haven’t looked through ALL of the steamer projections yet, but my two favorite projections so far:
McCann project to hit the most homers of his career
&
Yusmeiro Petit striking out 198 batters!
Whaa whaaat?!
You made me look them up! The McCann projection is reasonable, depending on what Steamer thinks the park effects should be in new Yankee Stadium. The Petit projection is definitely wrong: he should not be projected to start 30 games and relief in 35 other games.
My favorite projection to question so far is the Bill James Handbook projection for Michael Pineda. Look it up whenever that set of projections hits Fangraphs.
Bill James projections are always entertaining. I will keep Pineda in mind!
McCann might have too many PA projected between C/1B (587 PA is a career high) – hence the high HR total.
I know there was an issue with Petit because he was included on the starter and reliever depth chart. So his strikeouts are high because his IP is too high (208 IP projected). We’ll get that fixed.
Good call on the PA for McCann! His rate stats seem to be in line with his career more or less.
Is it really that ridiculous?
He’s got a career K/9 of 8.05. Steamer has him projected slightly above that rate at 8.58, which is still comfortably below this rate over the past two years (9.8).
If you think it’s ridiculous to project him out to 200+ inning, then that’s fair, but the strikeout potential is certainly legit.
The problem are his IP last year. Or, put more simply, will he start all year?
The average year-to-year IP gain for pitchers with an xFIP- less than 100 is something like 11 innings. He threw 117.0 last year.
I do think it’s ridiculous to project him out to 200+ innings (most innings EVER). Even more so when you see how those innings are amassed via 30 starts and 35 relief appearances (as noted in above comments from other posters). I do like his K rate though and would expect it to stay >8.00.
Don’t get me wrong, I’d love to see Petit match these projections as I have him in a dynasty league. His name just sticks out when you sort steamer projections by strikeouts.
It’s not just K’s either…I’m not sure that 3.22 ERA is realistic. His best ERA EVER at the age of 30 seems like a strange projection. Projection systems usually seem to be conservative and this is a best case scenario IMHO.
Just a funny steamer stand out is all!
I’d love to see fangraphs do a projection challenge. Give 100 player projections and have people guess over/under on their most confident 25 or something. That way, people can see just how hard this is (like NFL line-setting).
Something similar was done last season at the all-star break. The FG article had us guess 2nd half production based on projection or first half performance. Probably was a Sullivan article.
This is a nice article but contains a bit of a cop out at the end. Has the entire effort to remove the “bet which players will outperform expectations” statistically really been fully exhausted? Seems to me we should be at the point of going to that second step, with every projection being:
1. Here’s his steamer (for example) expectation.
2. And here’s his confidence interval to exceed that expectation based on his [3 year declining K-BB] [lucky BABIP last year] (fill in whatever statistic that we can shake out from all the years of data now)…
Of course in the end it still ends up a crapshoot and up to the fantasy owner’s gut. But still seems worthwhile to strive to have both an expectation (projection) and a statistical basis for betting the over/under right there alongside the projection.
An example of the problem is shown in the two recent articles discussing Andrew Cashner. One expert says Cashner won’t hit expectations based on 3 metrics, the other expert says he will exceed expectations based on 3 different metrics. There must be some way to calculate which of those numbers historically have been the most trustworthy. Otherwise we’re just left reading hours of articles (and not having the time to read hours more on every player) that sometimes contradict each other and a complete random choice what to believe in the end.
Those sorts of things are sort of all baked into steamer already AFAIK. Things like BABIP, HR/FB, etc are regressed toward the mean, until a player starts consistently beating the mean levels.
All players are regressed toward the mean, even if they’ve beaten it for 2,000 innings. It’s just that you regress less as you get more data. Also, I imagine that some peripheral stats (BABIP, HR/FB, etc.) are regressed more than others, but I’m not sure about that.
Play around with this for a while:
http://www.baseballheatmaps.com/graph/batter_marcel.php
It takes the hitters projected Marcel and shows how players with similar projections have performed have historically performed.
Now how about the same commentary for ZiPS! ZiPS seems to go more aggressively and projects a larger # of SPs under an ERA of 3 every year, and yet people compare the ZiPS/steamer lines like apples to apples…
Yeah, I like ZiPS a lot more than Steamer. The numbers seem more reasonable to me.
My first guess is that ZiPS may apply more weight to the mode than Steamer. The more you allow the mean to influence the projection, the more conservative it will be and vice versa.
Wouldn’t more emphasis on the mean result in a more positive projection?
The simple answer is regression. Every player’s projected performance is regressed toward the mean (though the extent differs by player). So, the set of projections are always going to gather up around the mean, or tighten the distribution. And then, of course, you’re going to see a wider spread of outcomes after the season because randomness generates outliers.
This is like ZiPs with batting average. Last year it only projected like 0 or 1 players projections to have an over .300 batting average but projected that there would be like 10-15 players who would achieve that or something.
It seems like one ought to be able to create confidence intervals for these projections. The distribution of outcomes for different players is probably not homogenous. I don’t know the ins and outs of how Steamer comes up with its forecast, but I’d hope the method allows for actually calculating the spread of each individual’s projected outcomes. That would allow us to then simulate the season many times and calculate on average how many players we expect to have an outcome like a sub 3.00 ERA.
Well, if a projection, as I paraphrase of the author, acts as some mixture of the mean and the mode, then your last sentence is exactly what the projection represents.
But I think I interpret your question as trying to say: How likely? 60%? 20%? The confidence interval would depict some expected range of outcomes across a number of standard deviations from the mean, which would be helpful in calculating the probability of achieving a particular value for a statistic.
For an example of how this works, look at a PECOTA card.
Is this the year that Rotographs will put out a value calculator customizable for different league settings? That’d be sweet. Maybe you’d get access as part of FanGraphs+?
If I might chime in, the first commentator, Mathew, has it closest to any of the responses here. Mix that with what Zezeil just said and you are 2/3 the way there. These projection systems are only able to base their projections on stats already accrued. A lot of players, especially those with large sample sizes over many seasons, it is easy to see what their projections should be and how Steamer, or Marcel, or Zips gets to its conclusions. It is basically a regression model that finds a particular mix between mean and mode.
HOWEVER, what the projections cannot know are the non-statistical changes that a player makes. It couldn’t have known that Garret Richards, for example, would add a complete new pitch and turn into a different player. But if we knew that he did, from reading press releases, news, etc etc etc, we can make some generalizations to augment the projection.
Furthermore, as I stated above, the projections as a Mean/Mode balancing act take into account all accrued statistics at the MLB level. But what if you are Jason Kipnis and an oblique injury sapped all of your power and bat speed. That is going to have a negative effect on the mean/mode projection of his stats. But if I know he’s going to be actually healthy now (or can at least assume so) then maybe I can augment the projection more positively.
Furthermore, i find projections to be only useful for independent stats. By that I mean stats that are more or less only affected by the player in questions. That means no runs or RBIs. Those are so dependent on what the players around you do. For example, Josh Donaldson (in my opinion) was a 95 RBI projection in Oakland. In the Toronto offense, however, I think he’s gonna be closer to a 100-105 RBI guy. Some projection systems take park factor into account (according to Donaldson he thinks he lost 14 HR to the spacious coliseum last year, direct quote). So that is another statistical category that requires manual augmentation as his HR/FB ration would have been adversely affected by his home park in Oakland.
Projections are great for showing what the pure, mathematical statistical regression looks like but requires significant manual augmentation based on opinion, news, and context in order to produce an actually accurate projection for a player prior to the season.
As a follow up, one of the new blurbs i always pay attention to is when players move teams and what pitching or batting coaches they will suddenly be working with. Certain coaches seem to have magic in their hands… for example it seemed like Dave Duncan had a way or turning every quasi-good SP into gold by teaching them how to play down in the zone and be willing to put balls into play as long as you were getting weak contact (versus always going for the K and such). Knowing that a guy with that kind of tutelage should see a rise in GB% should be an idicator that BABIP might change, BAA may change, HR/FB and FB% will also change, thus completely changing the potential final line outcome as far as ERA/xFIP/SIERRA go.
The only problem with adjusting a projection because Johnny Fastball added a pitch and we know he’s going to improve is that this kind of info looks much better in retrospect than it does before the fact. After Richards breaks out, we say “He added a new pitch – we should have predicted this.” Unfortunately there is far more noise than signal when it comes to “added a pitch”, “best shape of his life”, “changed his approach”, or whatever your favorite fundamental change would be. Every year was the year Bonderman was going to add a changeup, improve his changeup, try a splitter, spend spring training developing feel for his new changeup grip, or start throwing his changeup more. In simplest terms – most of what we know is wrong.
That’s a fair point. Obviously hindsight is awlays 20/20. But, if the goal is to produce projections, prior to a season, entering this information (such as a new pitching, or new pitching/hitting coach, conditioning program) should be included.
Of course, this can always lead you astray. I bet heavy on Kipnis last year 1) because I thought he could repeat elite stats mostly but 2) because I thought he could improve a lot based on news and commentary with the player about his off-season regimen. The result ended up being that he was hurt all year and thus, awful. But chances are he was hurt (oblique) because he was so bulked up that he had lost a lot of flexibility that he needed to be effective. So obviously it had the opposite effect. Nomar (during the late 90’s and early 2000’s had this problem… check out that SI front cover again if you don’t believe me)
Was just a thought to say that there are definite “non-statistical” elements that must be considered when making an accurate projection.
The statement about skewed distributions underneath the graph is in error. For all skewed right distributions, the mean is greater than the median. The median is a robust statistic and is not as influenced by the tail of the distribution. When a distribution is skewed LEFT, the mean is less than the median. The mean always “chases” the tail in comparison to the median. I hope this isn’t seen as snark. It was surprising to see such a fundamental error about non-symmetric distributions in print at FanGraphs. I think it was probably an honest typo.
He obviously meant to say left-skewed distribution, considering that’s what he drew. He indicated mean < median correctly if you assume such.
In the graph he sketched, which is skewed right, the mean would be bigger than the median. The mean get’s pulled toward the long tail.
Yikes, I should downvote myself. Regardless of the skew, yes, the mean should be closer to the tail than the median. Still giving the author the benefit of the doubt.
Yea I just transposed the two. I write at 6am, so mistakes happen. Not that it affects anything.
They also seem rather conservative on offense. I guess the model assumes players are more likely to under perform their true talent than over perform, due in part to injury and age related decline, so they get more accurate results at the individual level while at the group level it seems rather silly
I think you’ve got this all wrong. The model assumes that players will tend to perform near their true talent. The problem is that we don’t know what a player’s true talent is and thus have to estimate it using their individual observed performance and what we know about all players. Also, I think you have the second part backwards. The projection systems get very accurate results at the group level (think league-wide wOBA) and less so at the individual level due to randomness.
This is why by preferred projection system is PECOTA. Since the projection models are underspecified there is the opportunity for bringing other information sources to bear on their output. Because PECOTA gives a set of projections from 10 percent to 90 percent that allows the user, essentially, to move off of the median projection along the curve in either direction. Presenting only the median projection, on the other hand, provides a false level of confidence.