New Hitter xBABIP Based on BIS Batted Ball Data
You may have noticed that FanGraphs now feeds batted ball data, courtesy of Baseball Info Solutions, into its leaderboards. The day the data appeared, my mind buzzed with ways they could be useful in improving our understanding of a hitter’s batting average on balls in play (BABIP).
Mike Podhorzer already augmented previous attempts at devising an equation for expected batting average on balls in play (xBABIP) for hitters by incorporating elements of a hitter’s power, speed, plate discipline and batted ball tendencies. So, with fresh numbers in hand, I embarked on a journey to further improve the ever-evolving xBABIP. However, I sought to do so by using only batted ball data. Basically, I intended to develop a convenient xBABIP equation, one that can be computed using almost entirely variables found on the same page.
What I ultimately developed is a hitter xBABIP that is more a complement to Mike’s xBABIP than a substitute — that is, it is arguably no better than Mike’s equation, nor no worse, but it’s still different. I will explain what I mean in due time.
I chose batted ball variables that I thought would correlate well with BABIP (obviously):
- LD%, True FB% and True IFFB%: In his introduction to the new batted ball data, Tony Blengino demonstrated the frequencies with which certain types of batted balls turn into hits. Thus, optimal proportions of certain batted ball types could maximize BABIP. True IFFB% represents infield fly balls as a percentage of all balls in play, not just fly balls; it is calculated by multiplying IFFB% and FB%. True FB% denotes all fly balls minus infield flies. Econometric note: Because the fractions of every batted ball type sum to one (100 percent), I must omit one of them or else the regression will do it for me. That is why you do not see ground ball rate here.
- Hard%: Hard%, one of the new statistics, indicates how often a player hit a ball hard. (Who knew?) Eno Sarris was as surprised as I am to find that there is “virtually no correlation” between line drive rate and hard-hit percentage.
- Oppo%: Oppo%, another new statistic, indicates how often a player hit to the opposite field. One could argue in favor of using Pull%, the percentage of balls pulled, which would likely be negatively correlated with BABIP, especially for guys who encounter a ton of infield shifts. But righties experience shifts less often, so Pull% might not adequately capture the effect we might seek.
- Spd: I originally wanted to include infield hit rate (IFH%) so the equation would consist entirely of batted-ball variables. The idea was to capture (skilled) hits by bunts and (lucky) hits by dinkers and dribblers. However, per Jeff Zimmerman’s insight, I reconsidered my inclusion of them because they’re not necessarily exogenous to BABIP. A hitter’s speed score (Spd), on the other hand, is independent of infield hits; in other words, infield hits are a function of speed, not the other way around.
This model specification could be considered an expansion on Jeff’s work regarding hitter analytics in which he uses the aforementioned Hard% and Spd to generate expected BABIP values.
I limited the sample to all qualified hitters from 2002 through 2014, good for 1,971 observations. What follows are the results from the OLS regression:
xBABIP = .1975 — .4383*(True IFFB%) — .0914*(True FB%) + .2594*LD% + .1822*Hard% + .1198*Oppo% + .0042*Spd
Adjusted R-squared = .456
In light of Mike’s adjusted R-squared of .424, it’s clear that, on its own, more granular batted ball data hardly, let alone significantly, improve our understanding of BABIP. (I’ll interject and say it’s unwise to judge a model strictly by its R-squared, as there are a variety of statistical tests one can perform to test a model’s validity. But, alas, it is commonly used and more easily understood.) Even the improvement in year-to-year correlation is only the slightest upgrade to Mike’s results:
Y1 BABIP to Y2 BABIP: .4072
Y1 xBABIP to Y2 BABIP: .4712
So what can we conclude? For one, there isn’t necessarily a “correct” or “better” way to approach xBABIP — at least not yet. I’m sure if we threw the kitchen sink at the problem, everything would fall into place. But the sink would probably break and it would be really messy and no one would want to clean it up and that’s why we can’t have nice things. From what I observe, the new spray statistics (Pull%, Cent%, Oppo%) replace, rather than augment, absolute average angle, as used by Mike and provided by Baseball Heat Maps, and isolated power (ISO) serves as a proxy for the various degrees of contact quality as represented by Hard%, Med% and Soft%. Ultimately, it appears that despite having more precise batted ball data, we are not much closer to explaining away the luck component of BABIP in consideration of my attempt here — an attempt that is far from the be-all and end-all.
While my equation appears to be “better” at first glance, we won’t know for sure until the xBABIPs from Mike’s and my equations are compared side by side (in the form of, say, minimizing root mean squared error, or RMSE). Until then, indulge in the xBABIPs of 2015’s qualified hitters provided below. “Diff” represents the difference between xBABIP and BABIP; I conditionally formatted the cells so that blue indicates an overachiever and red an underachiever.

Alex, I’m not sure if it’s a problem on my end, but the Workbook isn’t loading. Error says it can’t be opened.
Heard that on Twitter — don’t know if that’s you or someone else. Trying to fix now. Sorry for the inconvenience!
Same here.
Nope, that was me. It’s still doing it, but again, thank you for taking a look into it!
I had the problem of it not being opened when my tab first opened, with there being an error noted, but the problem cleared up once I refreshed.
It could just be MS’s Azure cloud infrastructure, I tried to click to open it up into another tab, after refreshing, and it failed to do that too, the tab noted a service unavailablility error. “We are currently experiencing technical difficulties.
Please try again later.”
I eventually got it open, downloaded it, opened it in Google, then converted to Sheets format, in order to play around with it (using Chromebook).
Does Utley see a lot of defensive shifts?
Not sure about recently. He was shifted on about 9.9 percent of at-bats in 2013. I would guess that has increased.
http://www.fangraphs.com/fantasy/new-hitter-xbabip-based-on-bis-batted-ball-data/
The new soft/hard hits data is worthless without context of direction/angle hit. Plenty of “hard hit” balls by exit velocity off the bat are Outfield Flies that have very low probably BABIP.
There is a huge difference between a 95mph fly ball off the bat and a 95mph line drive. Unfortunately my understanding of the new hard/soft %’s are that they would bucket that 95mph off the bat lazy fly and that 95mph off the bat liner together making the “hard contact” stat not very useful for predicting babip.
It’s certainly not worthless, as in being without worth. It is obviously worth LESS than if we had perfect data, but it still has value.
We can break down the batted ball data not only by hardness but also by type. So we have 9 different classifications, really. GB, LD, FB and S, M, H. We can find an expected BABIP for each of those 9 possible batted ball types and come up with an expected BABIP after we project a payer’s future batted ball type distribution. That’s not useless, in fact that’s extremely useful. We can also find a spray chart for each batted ball type. So we can separate GBs by left side of the infield, and right side of the infield. We can use the spray charts to see how easy it would be to shift on a certain player. There’s all sorts of things we can do without having 100% perfect granular data.
If you re-read my comment I note that it’s worthless without context. You aptly describe ways to give it context. My point is that taking a raw “hard contact %” value that treats all 90+ mph exit velocity batted balls equally and expecting that value alone to help a babip prediction equation is foolish. Obviously other sorts of analyses can be done but as a babip predictor hard contact without further bucketing is not helpful.
Jeff Zimmerman has access to contact quality by batted ball type (he mentions it in his hitter analytics) but I’ve never found a public source. Might be private/proprietary. The search continues.
I also hesitate to assume the hard-hit data is blind to batted ball type. You have probably heard that HanRam had the highest velocity off the bat, but Jeff Sullivan debunked this. The new data didn’t claim that he had the highest Hard% nor the lowest Soft% — things one might assume based on his average batted ball velocity. Simply one anecdote, but perhaps a telling one.
My understanding from Appelman’s short review of the new stats, was that it classifies a hard FB differently than a hard liner, and then dumps both hard batted-ball types into hard%. But I think Scott’s point is just that there’s no way to tell from looking at hard% how many of those were FB or liners.
Right. Unfortunately, I share his same issue with the data — wish there was more context. Agh!
This is not correct. The soft/med/hard classified within the batted ball type. If the average line drive is 95mph and the average fly ball is 75 mph, a 95mph line drive would be classified as medium, and a 95 FB might be classified as hard.
Numbers for illustration purposes only.
And it looks like I was a few minutes late. Thanks for the clarification David, and I’m glad I had understood it correctly!
This makes a ton of sense. Thank you for clarifying. I am glad that my wrong interpretation was indeed that 🙂
FYI Scott, Skin Blues and anyone else in this mini-thread: each player has the fully contextual splits by batted ball and contact quality. Click the “Splits” tab, then hover over “Batted Balls” to choose a BIP type. Third of three tables is “Batted Ball,” which has not only contact quality but also ball spray percentages.
Example: http://www.fangraphs.com/statsplits.aspx?playerid=13611&position=OF&season=0&split=3.1
Yeah I’ve had a look at the splits, and it’s pretty awesome that the data is that granular. I can only imagine the amount of strain I put on your servers. Sorry!
On the topic of server strain, is it possible that in the future we will be able to do “Split Seasons” when looking at a range of years and simultaneously look at split stats? I attempted to look at K-BB% with batters on vs bases empty, and it won’t let me show data for a range of seasons split by year. That’s probably rough on the servers though and I may need to buy a second FG+ subscription to justify that feature…
That is fantastic. Thank you very much for explaining how to dig into that level of the data.
Did you use data from Year X in a regression to create the xBABIP equation, and then use that xBABIP equation to predict BABIP in Year X? In other words, did you use 2014 batted ball data to predict 2014 batted ball data? This is what Podhorzer did, and it explains the high correlation. In order to truly test it, you need to test the equation on out-of-sample data. Meaning, run a regression on the years 2002 to 2013, find your xBABIP equation, and then test the correlation with 2014 batted ball data. And if you want you can do the same thing for all other years in the sample. Then and only then can you truly compare the correlation of Y1 BABIP and Y1 xBABIP to Y2 BABIP. Of course xBABIP has a better correlation here, because the formula already knows the exact results of the 2014 season, which is an advantage that 2013 BABIP clearly does not have. With a 13 season sample size it won’t have as big of an effect as if you only used say 3 seasons of data since the bias is over 4 times smaller, but it will still have an effect on the correlation.
This is a good point. Oversight on my part.
I’m in a bit of a hurry but I quickly crunched this, using 2002-13 data and testing it out-of-sample on 2014:
2013 BABIP to 2014 BABIP: .4798
2013 xBABIP to 2014 BABIP: .4971
Will test on other years in sample.
Cool. Looks like it’s still better than BABIP, just by a smaller margin. I wonder if adding batter handedness as a variable would help it, by giving a boost to LHB. Also might be interesting to look at ground ball distribution. A lot of spread, and/or hitting the ball to the left of 2B, would probably correlate to a higher BA.
I tried creating a variable that measured the variance of a hitter’s batted ball spray, so that zero (or one) indicated a perfect 33-33-33 split by part of field and a larger number indicated a more predictable BIP pattern; it was mildly successful (over Oppo%) but needs tweaking so I omitted it.
The Y2Y correlation for 2014 was already strangely high so I’m skeptical about that particular one-year sample. But I’ll follow up with you with per-year updates.
now integrate shift effect!
I wish this stuff was posted on the main blog and not roto. I don’t play fantasy so I usually just skip the roto stuff due to a limited amount of time to read these articles, but occasionally they throw a gem like this on roto and I almost miss it. Anyways, awesome work Alex!
I appreciate the kind words! I’m (almost) strictly on the RG side — I could probably better coordinate with Dave Cameron in the future on things like these. Until then, keep your eyes peeled!
No worries, I’m sure for everyone like me who skips the fantasy stuff, there’s people who only read roto and not the main site. Glad I caught this piece though!
Hey Alex, good work. I think your model is probably suffering from omitted variable bias however. Why only include Hard%? Certainly soft and medium contact effect BABIP too, just not to the effect of Hard (presumably). Additionally, if flyballs and infield flyballs have an effect so should groundballs. I understand the problem with including variables that add to one, but you can get past this by using ratios as variables. Try using GB/FB as a variable or soft contact/medium contact. And what is the point of Oppo% alone? Only to control for susceptibility to the shift? I think omitting up the middle or pull is a mistake. How about you create a variable that measures the total variation in a batter’s batted ball locations? It could be as simple as finding the standard deviation of their ball locations (a 200 to center, 200 to pull, 200 to oppo player would be have a standard deviation of 0 and would be the ultimate spray hitter). Likewise, why not use more than a year’s worth of data? It wouldn’t be difficult to run a pooled cross section regression on all the years we have Hard% for.
I’ll try to address each question sequentially:
On their own, Hard% and Med% (with Soft% omitted) are statistically significant, but with the inclusion of other variables, only Hard% retains statistical significance. The problem is Hard% and Med% are fairly strongly correlated.
I tried ratios for BIP type; the estimates were less efficient.
I use Oppo% for simplicity. I mentioned in an earlier comment that I actually created a variable that captured a batter’s ability to spray the ball to all fields by calculating the variance of the three percentages. Replacing Oppo% with this variable is a microscopic improvement. But at that point, for that much work for such a small marginal gain, I opted for Oppo% for simplicity, as it had the strongest positive relationship. (Pull% and Oppo% are highly negatively correlated, so that their substitution affects little but the constant term.)
I use data from 2002 through 2014, representing all available data in terms of duration.
Let me know if you have any other questions!
Wow, Brandon Crawford has been hitting pretty well, but according to this, his BABIP should be much higher!
Kinda shocked to see Posey’s weak hitting, but that’s reflected in his drop in power so far this season, his poor hitting in cleanup has been a large part of the Giants offensive struggles.
Not too surprised to see Aoki and Panik doing well, right around what they have hit before for BABIP.
“I’m sure if we threw the kitchen sink at the problem, everything would fall into place. But the sink would probably break and it would be really messy and no one would want to clean it up and that’s why we can’t have nice things.”
Nailed it.
can you use integrals and derivatives to find Bapip