The Keys to Pitcher BABIP and HR/FB, Perhaps
Long has the relationship between pitcher performance and batted ball metrics been dubious. The Sabermetric community has a solid understanding of why, fundamentally, a pitcher is good or bad. Strikeouts are good. Walks are bad. Hits by pitch are also bad. Home runs allowed are especially bad. So on, so forth. And by no means are batted ball metrics useless. It’s how we know ground balls allowed are superior to fly balls allowed, for example.
The community had hoped, however, that more granular batted ball metrics would help us better explain some of the more nuanced elements of pitcher performance, including those related to luck, such as batting average on balls in play (BABIP) and the percentage of home runs per fly ball (HR/FB). Since their introduction to the public sphere in 2015, and even with the inclusion of more granular Statcast data in 2016, any relationships that might exist between the physics and outcomes for batted balls during an individual pitcher’s season are still poorly explained. The following table depicts the correlations between pitcher BABIP and various batted ball metrics, sorted by the strength of the relationship (all qualified seasons, 2007-17, n = 898):
| Metric | BABIP |
|---|---|
| BABIP | 1.000 |
| LD% | 0.348 |
| FB% | -0.329 |
| PU% | -0.326 |
| OFFB% | -0.307 |
| IFFB% | -0.239 |
| GB% | 0.204 |
| IFH% | 0.183 |
| Hard% | 0.179 |
| Soft% | -0.157 |
| BUH% | 0.110 |
| HR/FB | 0.083 |
| Med% | -0.033 |
| Cent% | 0.019 |
| Pull% | -0.017 |
| Oppo% | 0.004 |
OFFB% = outfield fly balls = FB% * (1 – IFFB%)
A coefficient of 1.0 indicates a perfect relationship (BABIP perfectly correlates with itself), and -1.0 indicates a perfect inverse relationship. There are some adequate correlations here — all of the batted ball types except for ground ball rate (GB%) exhibit weak correlations — but each of them, individually, explain little more than 10% of the variance in pitcher BABIP. Regressing BABIP on all of the batted ball types (LD%, GB%, OFFB%, PU%) produces an adjusted r2 of 0.22 — commendable, but it leaves unexplained almost 80% of the variance that a pitcher’s BABIP experiences in a given season. Meanwhile, none of the contact quality metrics (Hard%, Med%, Soft%) exhibit any semblance of a relationship with BABIP, and regressing BABIP against all three hardly inspires confidence (adjusted r2 = 0.14).
The thing is, we know, intuitively, in our hearts, that contact quality matters. The harder a ball is hit, the less time a fielder has to react to and cleanly field it; weak contact suggests the inverse. These relationships play themselves out for hitters more readily, affirming our understanding of the sport. But the lack of similar evidence in support of pitcher contact management is maddening.
Alas, our artificial truncation of per-observation sample sizes — anywhere from 162 innings to 253 innings (CC Sabathia, 2008) — theretofore truncates our understanding of these relationships. What happens if we expand sample sizes for each pitcher to, say, 500 innings? It may not be immediately helpful from a fantasy perspective (except to those who play in insane three-year leagues, perhaps), but lengthening the window during which each pitcher “collects” batted ball data significantly sharpens our correlation estimates (n = 287):
| Metric | 162+ IP | 500+ IP |
|---|---|---|
| BABIP | 1.000 | 1.000 |
| PU% | -0.326 | -0.531 |
| IFFB% | -0.239 | -0.503 |
| FB% | -0.329 | -0.430 |
| OFFB% | -0.307 | -0.395 |
| Soft% | -0.157 | -0.391 |
| GB% | 0.204 | 0.353 |
| LD% | 0.348 | 0.320 |
| Med% | -0.033 | 0.238 |
| Pull% | -0.017 | -0.180 |
| Cent% | 0.019 | 0.173 |
| HR/FB | 0.083 | 0.138 |
| IFH% | 0.183 | 0.137 |
| Hard% | 0.179 | 0.122 |
| Oppo% | 0.004 | 0.114 |
| BUH% | 0.110 | 0.035 |
OFFB% = outfield fly balls = FB% * (1 – IFFB%)
The batted ball types exhibit stronger correlations. Pop-ups, which are effectively automatic outs, explain almost 30% of the variance by themselves — a legitimately moderate correlation. Moreover, contact quality — namely soft contact (Soft%), and not hard contact (Hard%) — has borne what amounts to a non-zero correlation with BABIP. Lengthening the per-pitcher duration to 750 innings sharpens our correlations even further (n = 154):
| Metric | 162+ IP | 500+ IP | 750+ IP |
|---|---|---|---|
| BABIP | 1.000 | 1.000 | 1.000 |
| PU% | -0.326 | -0.531 | -0.582 |
| IFFB% | -0.239 | -0.503 | -0.549 |
| FB% | -0.329 | -0.430 | -0.485 |
| OFFB% | -0.307 | -0.395 | -0.450 |
| Soft% | -0.157 | -0.391 | -0.447 |
| GB% | 0.204 | 0.353 | 0.425 |
| Med% | -0.033 | 0.238 | 0.368 |
| LD% | 0.348 | 0.320 | 0.215 |
| Cent% | 0.019 | 0.173 | 0.213 |
| Pull% | -0.017 | -0.180 | -0.161 |
| IFH% | 0.183 | 0.137 | 0.120 |
| HR/FB | 0.083 | 0.138 | 0.099 |
| Oppo% | 0.004 | 0.114 | 0.070 |
| Hard% | 0.179 | 0.122 | 0.022 |
| BUH% | 0.110 | 0.035 | -0.009 |
OFFB% = outfield fly balls = FB% * (1 – IFFB%)
Note: the columns in all these tables are sortable!
Pop-ups now explain a third of BABIP variance on their own. But look! Soft and medium (Med%) contact demonstrate better-than-weak correlations, and the relationship between hard contact and BABIP is essentially nonexistent. This result is not to be conflated with the conclusion that hard contact doesn’t matter (don’t worry, it does) but it’s almost worthless as it relates to pitcher BABIP, whereas its inverse companions are much more valuable. We must acknowledge the significantly smaller sample size here — 154 player-seasons is still fairly small — but it’s also self-evident these kinds of relationships, while definitely existing, don’t necessarily flesh themselves out over the course of one season.
Using the sample of 750+ inning observations, regressions of BABIP on batted ball types, contact quality, and directions bear the following correlations (as measured by adjusted r2):
LD%, GB%, OFFB%, PU%: 0.40
Soft%, Med%, Hard%: 0.23
Pull%, Cent%, Oppo%: 0.03 (this likely bears a stronger correlation from the hitter side, especially when controlling for hitter handedness)
The takeaway? Most everything we’ve known, or, through intuition, thought we’ve known, about baseball is probably correct in the grand scheme of things. It’s just that we may need to look at it through a slightly different lens. Soft and medium contact affect BABIP much more significantly than hard contact does. It helps explain why guys like Dallas Keuchel, Jake Arrieta, Kyle Hendricks, Tanner Roark, and Marco Estrada have a habit of running impressively low BABIPs — and also, in a roundabout way, explains why the jury is still out on Robbie Ray’s insanely volatile annual BABIPs. Maybe now looking at this year’s BABIP leaders, sorted by Soft%, won’t surprise you.
HR/FB
HR/FB, another pitcher metric deeply embedded with luck, also improves in correlation/explanation as pitcher sample sizes increase. Using all the aforementioned samples:
| Metric | 162+ IP | 500+ IP | 750+ IP |
|---|---|---|---|
| HR/FB | 1.000 | 1.000 | 1.000 |
| Oppo% | -0.227 | -0.418 | -0.431 |
| PU% | -0.169 | -0.283 | -0.377 |
| IFFB% | -0.166 | -0.280 | -0.376 |
| Pull% | 0.255 | 0.369 | 0.376 |
| GB% | 0.111 | 0.229 | 0.346 |
| FB% | -0.139 | -0.248 | -0.342 |
| OFFB% | -0.122 | -0.233 | -0.326 |
| Soft% | -0.143 | -0.265 | -0.261 |
| Hard% | 0.374 | 0.362 | 0.240 |
| BUH% | 0.059 | 0.234 | 0.177 |
| LD% | 0.073 | 0.050 | -0.099 |
| BABIP | 0.083 | 0.138 | 0.099 |
| Cent% | -0.086 | -0.029 | -0.032 |
| IFH% | 0.023 | -0.005 | 0.025 |
| Med% | -0.237 | -0.140 | -0.019 |
OFFB% = outfield fly balls = FB% * (1 – IFFB%)
Pitcher HR/FB behaves so much differently than pitcher BABIP. For example, batted ball direction plays a pretty huge role: pulled batted balls are bad, and oppo batted balls are good. That makes sense; on average, hitters hit for twice as much power to their pull side than to the opposite field. There’s some weirdness in the table that suggests the correlation between hard-hit rate and HR/FB dramatically decreases as pitchers accumulate more innings, but I’m convinced it’s nothing more than a data quirk: when I limited the data to pitcher-seasons 1,000+ innings (n = 99), the correlation coefficient increases to 0.406.
So, hard contact is bad, of course. And so are pulled balls, especially of the fly ball variety. Lastly, a minor interesting note (something Jeff Zimmerman has previously shown, I think): ground ball pitchers tend to allow higher HR/FB rates.
Again, I’m not sure this is anything new to us. This, for HR/FB and for BABIP, is all intuition. These are relationships that some simple manipulation of FanGraphs’ splits leaderboard would reveal to us. But there’s something about seeing these correlations emerge over time at the individual pitcher level that is comforting, reassuring. Sure, your favorite pitcher might still be subject to a ton of good or bad luck in any given season. But there’s also evidence that there’s a method to this madness. There’s evidence there still exists, in some form or another, The ProcessTM.
It genuinely weirds me out that the correlations with Hard% decrease as data size increases. Especially for HRs. It seems like we should all be banned from using Hard% until we have a satisfying explanation of that.
There’s a note immediately following that table wherein I posit the “decreasing” correlation with HR/FB is a data quirk. At 1,000+ IP, the relationship between Hard% and HR/FB is at its most robust.
That is somewhat reassuring, thanks. But it still feels a bit like we’re cherrypicking an endpoint — do we want to peek at whether correlation bumps down again at 2000?
And even 500 is a lot of innings. If the data is quirky that far out… I don’t know, I’m not a statistician, it just seems like a warning.
This might shock you (I am being completely sincere), but only eight pitchers have accumulated 2,000+ innings since 2007. It’s just not a feasible endpoint. That said, you kind of have to cherry-pick your endpoint. I could’ve done it for every 100 (or even every 50) innings, but my intention wasn’t to be incredibly precise. I wanted to know if relationships exist long-term; they do, and they tend to get stronger, which is generally everything we’ve come to know about the reliability of metrics over certain sample sizes. It’s just that it takes much longer for these things to become reliable for pitchers than they do hitters, so much so that pitchers are subject to a large amount of variance within a given season. But they even out over time.
500 innings is a lot of innings, but it might not be as many as we think. I think that’s kind of the point: 200 feels like a lot of innings, yet it’s not nearly enough to even begin to bear meaningful relationships. So 500, which feels like a ton, might only start to be an adequate amount of time to allow these relationships to flesh out. (It could be sooner — as aforementioned, I could be more systematic about testing more frequent innings increments — but I simply wanted to cast my line farther out into the water.)
Thoughts on the inherent blind spot of selection bias in cumulative innings analysis, especially in HR/FB? For instance, a pitcher that allows a 20%+ HR/FB isn’t going to get 500 IP in the majors, so all of their data is filtered out of the bigger samples. Meaning, generally, effective pitchers are overrepresented in those large IP samples.
Absolutely true, although there are some pretty insanely bad pitchers that have accumulated 500 innings. They always find their way onto roster somehow.
But you’re generally right. If you look at, for example, BABIP among all full-time hitters and pitchers, they will inevitably be higher and lower (respectively) than league-average due to some combination of skill and luck. The same would apply to HR/FB suppression for pitchers. I’m just not sure it plays a massive role.
Rambling tangent: pop flies are something we know has a strong relation to BABIP, in the sense that they’re almost always caught. When we see limited correlation same-year, that’s not telling about the pop flies themselves, it’s just saying they’re a limited part of the whole picture. I wonder, is there a statistical tool that would make this clearer?
“all of the batted ball types except for ground ball rate (GB%) exhibit weak correlations“
I’m probably being dense here, but doesn’t LD% have a stronger correlation (.348) than GB% (.204)?
Yeah — I should clarify: all of them, except for GB%, exhibit some correlation; it’s just that they’re all weak correlations. GB% = effectively no correlation. Sorry for the confusion!
Gotcha. Thanks for clarifying!
Awesome article, Alex. Do you think you’ve finally solved Podhorzer’s 2015 “Brandon Mccarthy Challenge” to prove his high hr/fb isn’t just bad luck?
It seems like a good next step would be to test the predictive validity of things like hard contact and iffb rate on future hr/fb rate.
https://www.fangraphs.com/fantasy/challenge-prove-brandon-mccarthys-hrfb-is-not-bad-luck/
This is a difficult question to answer only because there’s no distinction between HR/FB as a descriptor or a predictor. Descriptively, maybe McCarthy did deserve those home runs. But player performance ebbs and flows, and predictively, there’s almost no reason to believe his HR/FB would remain that bad.
HR/FB is so heavily afflicted by variance that even though there appears to be some relationship between HR/FB and, say, Hard%, I’m reluctant to even trust that relationship to flesh out in a given season. (I’m also not sure how well Hard% sticks year over year, but maybe that’s beside the point.) It’s difficult to reconcile; it’ll take some mental re-training on my part to trust that something like Hard% actually does affect HR/FB. It just might not bear fruit immediately — and, moreover, the Hard% will fluctuate over time — making it difficult to fully trust as an in-season indicator of good or bad luck. (I mean, sometimes there’s clearly good or bad luck, like in McCarthy’s case, where we’re describing “bad luck” as “very likely to regress toward league average during the remainder of the season, if he were to pitch” (which he did not). But identifying the extreme outliers is always easier than, say, trying to figure out if Chris Archer is actually just slightly below average at allowing home runs per fly ball.)
Wow, long comment, sorry!!!
Quite insightful, thanks for elaborating! It sounds like any sustainable talent in suppressing (or allowing extra) homeruns per flyball is difficult to convincingly prove…and probably pretty modest in effect size.
Alex – as always, great and interesting work. 2 quick questions:
1)for BABIP, if you remove pop-ups from soft contact, does soft contact still matter or is it pop-ups that matter and soft-contact is picking that up so that a metric of (soft contact less pop-ups) would have a much weaker correlation?
2) For HR/FB, as pulled FB seem to be an important factor is there a regression back to the mean for pitchers from one year to the next or are some pitchers prone to pulled FB and likely then (holding everything else equal) a higher HR/FB rate?
1) You’re likely right that Soft% is picking up IFFB%, whether partially or entirely. I wish I could tease out the effect; the splits leaderboard doesn’t allow cross-cutting by IFFB%. Regardless I think you’ve generally made the right assertion, one that I overlooked! I’m clearly a bit rusty.
2) I recall testing the year-over-year relationship of pulled fly balls at one point for xISO but cannot recall the specifics. It’s definitely worth revisiting. Curious how it might bear a relationship to fastball velocity.
Note: Sortable columns in the tables is freaking amazing! Can we somehow standardize this across FG articles?
Doesn’t seem possible that OFFB% has a negative corr with HR/FB given that its just the inverse of IFFB% which has a strong negative corr. Oh yeah, it also makes no sense.
Great concept and article.