Projecting the Impossible: Pitcher Wins
In the latest episode of the Launch Angle Podcast, Rob Silver asked me how many Wins did I expect Chris Archer to accumulate this season. Basically, I came back with my normal response, I don’t chase Wins and don’t care. He pushed a little harder and wondered the actual difference. I just stammered out a horrible response because I didn’t know. I’m not one to not know so found out with the answer being a win or two.
For years, I’ve used the potential for more Wins as a tie breaker between pitchers with similar baseline stats (strikeouts, walks, and groundball rate). I focused on talent first. Usually, I found pitchers on projected better teams being drafted way ahead of those with similar skills on worse teams. I just assumed the better skills will lead the pitcher to as many Wins as the worse pitcher on a better team. There is no need for me to make that assumption anymore.
The data sources are now available to see how many Wins a pitcher is projected to win knowing his projected ERA and his team’s projected win percentage. For the projections, I used Steamer projections. For the preseason projected winning percentage for each team, I used the CAIRO projections from the Replacement Level Yankees Weblog from 2010 to 2015 and our FanGraphs projected standing from the past two seasons.
Then, I matched up projected starters (GS >= 20, GS/G >= .95) and the teams with the actual results. I dropped the requirements down to 10 GS and GS/G to 50% to account for injuries and teams not giving bad starters many starts. In all, I matched up 861 starters.
The first test I ran was a linear regression with the projected ERA and team win% to the pitcher’s actual Wins per games started. The r-squared ended up at 0.15 (0.11 with just ERA). Not good or surprising considering all the factors involved. The equation works out to:
Actual Pitcher Win% = .282-0.0585*proj ERA+.646*proj team win%
| Team Win%/ERA | 3.00 | 3.50 | 4.00 | 4.50 | 5.00 |
|---|---|---|---|---|---|
| .400 | 29.4% | 29.3% | 29.2% | 29.1% | 29.0% |
| .425 | 31.2% | 31.1% | 31.0% | 30.9% | 30.8% |
| .450 | 33.0% | 32.9% | 32.8% | 32.7% | 32.7% |
| .475 | 34.8% | 34.7% | 34.7% | 34.6% | 34.5% |
| .500 | 36.7% | 36.6% | 36.5% | 36.4% | 36.3% |
| .525 | 38.5% | 38.4% | 38.3% | 38.2% | 38.1% |
| .550 | 40.3% | 40.2% | 40.1% | 40.0% | 39.9% |
| .575 | 42.1% | 42.0% | 41.9% | 41.8% | 41.7% |
| .600 | 43.9% | 43.8% | 43.7% | 43.6% | 43.6% |
While these values give the Win%, it’s not an easy number to work with. Instead, here are the Win totals given 32 starts.
| Team Win%/ERA | 3.00 | 3.50 | 4.00 | 4.50 | 5.00 |
|---|---|---|---|---|---|
| .400 | 9.4 | 9.4 | 9.3 | 9.3 | 9.3 |
| .425 | 10.0 | 10.0 | 9.9 | 9.9 | 9.9 |
| .450 | 10.6 | 10.5 | 10.5 | 10.5 | 10.5 |
| .475 | 11.1 | 11.1 | 11.1 | 11.1 | 11.0 |
| .500 | 11.7 | 11.7 | 11.7 | 11.6 | 11.6 |
| .525 | 12.3 | 12.3 | 12.3 | 12.2 | 12.2 |
| .550 | 12.9 | 12.9 | 12.8 | 12.8 | 12.8 |
| .575 | 13.5 | 13.4 | 13.4 | 13.4 | 13.4 |
| .600 | 14.1 | 14.0 | 14.0 | 14.0 | 13.9 |
Now, there is quite a bit of difference for teams on the extreme ends with about 5 Wins separating the same pitcher on a team projected with a .600 winning percentage vice one at .400.
Using linear regression, the team winning percentage is more of a factor than talent in the number of Wins a pitcher gets.
Besides linear regression, I just bucketed the data in 0.25 ERA blocks with either a projected winning or losing record.
First, here are the average and median values for each pitcher block.
Average Results
| Winning% | Wins w/ 32 GS | |||
|---|---|---|---|---|
| ERA | Winning Team | Losing Team | Winning Team | Losing Team |
| 3.00 or less | 54% | 53% | 17.2 | 16.9 |
| 3.01 to 3.25 | 45% | 40% | 14.3 | 12.8 |
| 3.26 to 3.50 | 44% | 38% | 14.2 | 12.1 |
| 3.51 to 3.75 | 42% | 38% | 13.3 | 12.0 |
| 3.76 to 4.00 | 39% | 34% | 12.5 | 10.7 |
| 4.01 to 4.25 | 38% | 33% | 12.0 | 10.6 |
| 4.25 to 4.50 | 38% | 32% | 12.2 | 10.3 |
| 4.51 to 4.75 | 35% | 32% | 11.2 | 10.3 |
| 4.76 to 5.00 | 30% | 30% | 9.7 | 9.7 |
| 5.01 or more | 36% | 29% | 11.6 | 9.1 |
Median Results
| Winning% | Wins w/ 32 GS | |||
|---|---|---|---|---|
| ERA | Winning Team | Losing Team | Winning Team | Losing Team |
| 3.00 or less | 57% | 53% | 18.1 | 17.0 |
| 3.01 to 3.25 | 44% | 42% | 14.1 | 13.3 |
| 3.26 to 3.50 | 45% | 37% | 14.4 | 11.8 |
| 3.51 to 3.75 | 40% | 37% | 12.8 | 11.7 |
| 3.76 to 4.00 | 39% | 33% | 12.6 | 10.7 |
| 4.01 to 4.25 | 38% | 32% | 12.2 | 10.3 |
| 4.25 to 4.50 | 38% | 30% | 12.0 | 9.7 |
| 4.51 to 4.75 | 36% | 32% | 11.6 | 10.4 |
| 4.76 to 5.00 | 28% | 31% | 9.0 | 9.9 |
| 5.01 or more | 35% | 29% | 11.2 | 9.1 |
The difference is around 1.5 Wins between being on a winning or losing team. This difference is about the same as the difference between .475 and .525 teams in the linear regression model.
The value I find most interesting is that around a pitcher with a 4.50 projected ERA on a projected winning team will get as many Wins as a 3.50 projected ERA pitcher on a projected losing team.
Let’s use this information to examine Archer. He’s projected for several ERA’s (3.27 to 3.58) so I’ll assume a 3.50 ERA for this example. The Rays are projected for a winning%, at .486. Here are his projected Wins from his projections and the above process.
Source: Wins (per 32 GS), Win%
Averaged projections: 13.8, 43%
Linear regression: 11.4, 36%
Average binned: 12.0, 38%
Median binned: 11.8, 37%
Hold while I try one more method before I compare the results. I took the pitchers with a projected ERA from 3.25 to 3.75 with a team projected win% between .456 to .513. In all, 23 pitchers were in the sample and averaged a 41.1% Win% (median was 40.6%) with the Wins working out to 13.2 for the average value (median was 13.0).
It’s a range from 11 to 14 Wins for Archer. If on a projected winning team, he’d be projected between 12 and 16 wins. There is a lot of projection going on but that’s all we can really do at this point in the season.
Overall, I expected the Win total difference to be smaller depending on the if the team is projected to have a winning or losing record. The difference works out to about 1.5 Wins and can grow or shrink a bit depending on the team’s projected win%. For me, I’d be looking at the extremes (e.g. Astros, Royals, Marlins, and White Sox) to find those pitchers who could break the norm. Most teams are near the middle so we need to project an average number of Wins from them.
Jeff, one of the authors of the fantasy baseball guide,The Process, writes for RotoGraphs, The Hardball Times, Rotowire, Baseball America, and BaseballHQ. He has been nominated for two SABR Analytics Research Award for Contemporary Analysis and won it in 2013 in tandem with Bill Petti. He has won four FSWA Awards including on for his Mining the News series. He's won Tout Wars three times, LABR twice, and got his first NFBC Main Event win in 2021. Follow him on Twitter @jeffwzimmerman.
What sort of error bars are there on these?
It’s interesting that, at least for the regression, perhaps team winning percentage is acting as a somewhat proxy for team defense and/or offense.
> Actual Pitcher Win% = .282-0.0585*proj ERA+.646*proj team win%
Great, my Gerrit Cole + Charlie Morton should have 40 wins upside.
Could be. I wanted to keep it simple for now. Tons of factors go into it but as Rob did, most people just take into account expectations.
I think this is what you’re highlighting, but – the biggest factor is IP. Both GS and IP/GS are the overwhelming controlling factors in pitcher W’s.
This will become more important as time goes on as well. There is going to be some interaction between talent and IP/GS that will needed to be accounted for as more pitchers max out at 2 time through the order. This will skew ERA’s lower and possibly make the relationship to wins decrease lower than it already is.
In contrast, Brad Peacock won 10 games in 21 starts. So maybe it wont matter that much either.
As someone who had Rich Hill in a QS league, this is exactly it right here.
Sort of. Being good (low ERA) means the pitcher will throw more innings. Pitches per inning might be a better measure. It’s something I can write up later.
I did run a quick correlation between IP/GS and W/GS and the correlation was .22. It will be lower by using projected IP/GS
It’s probably more useful to use team runs scored than winning percentage as an input, as the quality of the other SPs on the team shouldn’t affect the answer.
I which. I have to include park factors with places like Colorado. I’m trying to keep it simple.
I wonder if you could use similar logic to see which teams middle relievers have the best chance of vulturing wins, ex: would Green with the Yankees be a better bet for wins than Miller for the Indians because of the winning ratio of the team and how deep the starters go into the game? the less innings from the starter, the better chance on the bullpen grabbing the goods (win)?
I’ve thought about this a lot more recently. Every year a rp goes in the 10 win range, and more and more are sneaking up that way.
I wonder if by reverse engineering the answer could find guys to target.
Throws 100 IP
Winning team
Good
Wouldn’t you get much, much closer if you also took into account average innings pitched per outing? I’d expect that the 5 innings guys are much less likely to benefit from a winning team than the 7 inning guys.
Good stuff. As part of my off-season pitcher analysis, looking for Win opp and risk is one of the final refinements I make. I’ve used Patrick Davitt (Baseball HQ) formula of eWins=((RS/G^1.8)/(ERA^1.8)+RS/G^1.8))*(.72*GS). I don’t have the math chops to validate it, but it has always yielded some intuitive insights. For example, compared to Steamer and BaseballHQ win projections this year, ID’d Thor, McCullers, Musgrove, Walker, and S Gray on upside (all at least 4 wins above projection), and McHugh, Eovaldi, Conley, Flaherty on downside (again, all at least 4 below projection). Like you, I use to break ties. You planning to look at non-linear models?
I haven’t looked into this model (and other recommended on Twitter).
Using a linear regression, the idea is that all variables act as “all else equal” impacts on the result (in this case, wins). I think it would be more accurate to calculate expected wins for the team in non-starts (i.e. total team wins less expected wins for that SP). To avoid a circular reference situation, maybe just subtract out the previous season’s team wins per that SP’s start and have that be the isolated team win rate.
Otherwise, you’re not isolating the two variables. A more elegant solution is probably out there, but in the current format, you are measuring two variables with a causal relationship between them.
Agree. It’s either a simple or crazy complicated solution.