New Hitter xBABIP Based on BIS Batted Ball Data

You may have noticed that FanGraphs now feeds batted ball data, courtesy of Baseball Info Solutions, into its leaderboards. The day the data appeared, my mind buzzed with ways they could be useful in improving our understanding of a hitter’s batting average on balls in play (BABIP).

Mike Podhorzer already augmented previous attempts at devising an equation for expected batting average on balls in play (xBABIP) for hitters by incorporating elements of a hitter’s power, speed, plate discipline and batted ball tendencies. So, with fresh numbers in hand, I embarked on a journey to further improve the ever-evolving xBABIP. However, I sought to do so by using only batted ball data. Basically, I intended to develop a convenient xBABIP equation, one that can be computed using almost entirely variables found on the same page.

What I ultimately developed is a hitter xBABIP that is more a complement to Mike’s xBABIP than a substitute — that is, it is arguably no better than Mike’s equation, nor no worse, but it’s still different. I will explain what I mean in due time.

I chose batted ball variables that I thought would correlate well with BABIP (obviously):

  • LD%, True FB% and True IFFB%: In his introduction to the new batted ball data, Tony Blengino demonstrated the frequencies with which certain types of batted balls turn into hits. Thus, optimal proportions of certain batted ball types could maximize BABIP. True IFFB% represents infield fly balls as a percentage of all balls in play, not just fly balls; it is calculated by multiplying IFFB% and FB%. True FB% denotes all fly balls minus infield flies. Econometric note: Because the fractions of every batted ball type sum to one (100 percent), I must omit one of them or else the regression will do it for me. That is why you do not see ground ball rate here.
  • Hard%: Hard%, one of the new statistics, indicates how often a player hit a ball hard. (Who knew?) Eno Sarris was as surprised as I am to find that there is “virtually no correlation” between line drive rate and hard-hit percentage.
  • Oppo%: Oppo%, another new statistic, indicates how often a player hit to the opposite field. One could argue in favor of using Pull%, the percentage of balls pulled, which would likely be negatively correlated with BABIP, especially for guys who encounter a ton of infield shifts. But righties experience shifts less often, so Pull% might not adequately capture the effect we might seek.
  • Spd: I originally wanted to include infield hit rate (IFH%) so the equation would consist entirely of batted-ball variables. The idea was to capture (skilled) hits by bunts and (lucky) hits by dinkers and dribblers. However, per Jeff Zimmerman’s insight, I reconsidered my inclusion of them because they’re not necessarily exogenous to BABIP. A hitter’s speed score (Spd), on the other hand, is independent of infield hits; in other words, infield hits are a function of speed, not the other way around.

This model specification could be considered an expansion on Jeff’s work regarding hitter analytics in which he uses the aforementioned Hard% and Spd to generate expected BABIP values.

You Aren't a FanGraphs Member
It looks like you aren't yet a FanGraphs Member (or aren't logged in). We aren't mad, just disappointed.
We get it. You want to read this article. But before we let you get back to it, we'd like to point out a few of the good reasons why you should become a Member.
1. Ad Free viewing! We won't bug you with this ad, or any other.
2. Unlimited articles! Non-Members only get to read 10 free articles a month. Members never get cut off.
3. Dark mode and Classic mode!
4. Custom player page dashboards! Choose the player cards you want, in the order you want them.
5. One-click data exports! Export our projections and leaderboards for your personal projects.
6. Remove the photos on the home page! (Honestly, this doesn't sound so great to us, but some people wanted it, and we like to give our Members what they want.)
7. Even more Steamer projections! We have handedness, percentile, and context neutral projections available for Members only.
8. Get FanGraphs Walk-Off, a customized year end review! Find out exactly how you used FanGraphs this year, and how that compares to other Members. Don't be a victim of FOMO.
9. A weekly mailbag column, exclusively for Members.
10. Help support FanGraphs and our entire staff! Our Members provide us with critical resources to improve the site and deliver new features!
We hope you'll consider a Membership today, for yourself or as a gift! And we realize this has been an awfully long sales pitch, so we've also removed all the other ads in this article. We didn't want to overdo it.

I limited the sample to all qualified hitters from 2002 through 2014, good for 1,971 observations. What follows are the results from the OLS regression:

BABIP vs xBABIP

xBABIP = .1975 — .4383*(True IFFB%) — .0914*(True FB%) + .2594*LD% + .1822*Hard% + .1198*Oppo% + .0042*Spd
Adjusted R-squared = .456

In light of Mike’s adjusted R-squared of .424, it’s clear that, on its own, more granular batted ball data hardly, let alone significantly, improve our understanding of BABIP. (I’ll interject and say it’s unwise to judge a model strictly by its R-squared, as there are a variety of statistical tests one can perform to test a model’s validity. But, alas, it is commonly used and more easily understood.) Even the improvement in year-to-year correlation is only the slightest upgrade to Mike’s results:

Y1 BABIP to Y2 BABIP: .4072
Y1 xBABIP to Y2 BABIP: .4712

So what can we conclude? For one, there isn’t necessarily a “correct” or “better” way to approach xBABIP — at least not yet. I’m sure if we threw the kitchen sink at the problem, everything would fall into place. But the sink would probably break and it would be really messy and no one would want to clean it up and that’s why we can’t have nice things. From what I observe, the new spray statistics (Pull%, Cent%, Oppo%) replace, rather than augment, absolute average angle, as used by Mike and provided by Baseball Heat Maps, and isolated power (ISO) serves as a proxy for the various degrees of contact quality as represented by Hard%, Med% and Soft%. Ultimately, it appears that despite having more precise batted ball data, we are not much closer to explaining away the luck component of BABIP in consideration of my attempt here — an attempt that is far from the be-all and end-all.

While my equation appears to be “better” at first glance, we won’t know for sure until the xBABIPs from Mike’s and my equations are compared side by side (in the form of, say, minimizing root mean squared error, or RMSE). Until then, indulge in the xBABIPs of 2015’s qualified hitters provided below. “Diff” represents the difference between xBABIP and BABIP; I conditionally formatted the cells so that blue indicates an overachiever and red an underachiever.





Two-time FSWA award winner, including 2018 Baseball Writer of the Year, and 8-time award finalist. Featured in Lindy's magazine (2018, 2019), Rotowire magazine (2021), and Baseball Prospectus (2022, 2023, 2024, 2025). Biased toward a nicely rolled baseball pant.

33 Comments
Oldest
Newest Most Voted
TangoAlphaLima
11 years ago

Alex, I’m not sure if it’s a problem on my end, but the Workbook isn’t loading. Error says it can’t be opened.

Jackie T.
11 years ago

Same here.

dj_mosfett
11 years ago

Nope, that was me. It’s still doing it, but again, thank you for taking a look into it!

obsessivegiantscompulsive
11 years ago

I had the problem of it not being opened when my tab first opened, with there being an error noted, but the problem cleared up once I refreshed.

It could just be MS’s Azure cloud infrastructure, I tried to click to open it up into another tab, after refreshing, and it failed to do that too, the tab noted a service unavailablility error. “We are currently experiencing technical difficulties.
Please try again later.”

I eventually got it open, downloaded it, opened it in Google, then converted to Sheets format, in order to play around with it (using Chromebook).

cdarcyMember since 2016
11 years ago

Does Utley see a lot of defensive shifts?

Scott
11 years ago

The new soft/hard hits data is worthless without context of direction/angle hit. Plenty of “hard hit” balls by exit velocity off the bat are Outfield Flies that have very low probably BABIP.

There is a huge difference between a 95mph fly ball off the bat and a 95mph line drive. Unfortunately my understanding of the new hard/soft %’s are that they would bucket that 95mph off the bat lazy fly and that 95mph off the bat liner together making the “hard contact” stat not very useful for predicting babip.

RotoholicMember since 2016
11 years ago
Reply to  Scott

It’s certainly not worthless, as in being without worth. It is obviously worth LESS than if we had perfect data, but it still has value.

We can break down the batted ball data not only by hardness but also by type. So we have 9 different classifications, really. GB, LD, FB and S, M, H. We can find an expected BABIP for each of those 9 possible batted ball types and come up with an expected BABIP after we project a payer’s future batted ball type distribution. That’s not useless, in fact that’s extremely useful. We can also find a spray chart for each batted ball type. So we can separate GBs by left side of the infield, and right side of the infield. We can use the spray charts to see how easy it would be to shift on a certain player. There’s all sorts of things we can do without having 100% perfect granular data.

Scott
11 years ago
Reply to  Rotoholic

If you re-read my comment I note that it’s worthless without context. You aptly describe ways to give it context. My point is that taking a raw “hard contact %” value that treats all 90+ mph exit velocity batted balls equally and expecting that value alone to help a babip prediction equation is foolish. Obviously other sorts of analyses can be done but as a babip predictor hard contact without further bucketing is not helpful.

wildcard09
11 years ago

My understanding from Appelman’s short review of the new stats, was that it classifies a hard FB differently than a hard liner, and then dumps both hard batted-ball types into hard%. But I think Scott’s point is just that there’s no way to tell from looking at hard% how many of those were FB or liners.

David AppelmanFanGraphs Staff
11 years ago
Reply to  Scott

This is not correct. The soft/med/hard classified within the batted ball type. If the average line drive is 95mph and the average fly ball is 75 mph, a 95mph line drive would be classified as medium, and a 95 FB might be classified as hard.

Numbers for illustration purposes only.

wildcard09
11 years ago
Reply to  David Appelman

And it looks like I was a few minutes late. Thanks for the clarification David, and I’m glad I had understood it correctly!

Scott
11 years ago
Reply to  David Appelman

This makes a ton of sense. Thank you for clarifying. I am glad that my wrong interpretation was indeed that 🙂

RotoholicMember since 2016
11 years ago

Yeah I’ve had a look at the splits, and it’s pretty awesome that the data is that granular. I can only imagine the amount of strain I put on your servers. Sorry!

On the topic of server strain, is it possible that in the future we will be able to do “Split Seasons” when looking at a range of years and simultaneously look at split stats? I attempted to look at K-BB% with batters on vs bases empty, and it won’t let me show data for a range of seasons split by year. That’s probably rough on the servers though and I may need to buy a second FG+ subscription to justify that feature…

Scott
11 years ago

That is fantastic. Thank you very much for explaining how to dig into that level of the data.

RotoholicMember since 2016
11 years ago

Did you use data from Year X in a regression to create the xBABIP equation, and then use that xBABIP equation to predict BABIP in Year X? In other words, did you use 2014 batted ball data to predict 2014 batted ball data? This is what Podhorzer did, and it explains the high correlation. In order to truly test it, you need to test the equation on out-of-sample data. Meaning, run a regression on the years 2002 to 2013, find your xBABIP equation, and then test the correlation with 2014 batted ball data. And if you want you can do the same thing for all other years in the sample. Then and only then can you truly compare the correlation of Y1 BABIP and Y1 xBABIP to Y2 BABIP. Of course xBABIP has a better correlation here, because the formula already knows the exact results of the 2014 season, which is an advantage that 2013 BABIP clearly does not have. With a 13 season sample size it won’t have as big of an effect as if you only used say 3 seasons of data since the bias is over 4 times smaller, but it will still have an effect on the correlation.

RotoholicMember since 2016
11 years ago

Cool. Looks like it’s still better than BABIP, just by a smaller margin. I wonder if adding batter handedness as a variable would help it, by giving a boost to LHB. Also might be interesting to look at ground ball distribution. A lot of spread, and/or hitting the ball to the left of 2B, would probably correlate to a higher BA.

rotobanter
11 years ago

now integrate shift effect!

wildcard09
11 years ago

I wish this stuff was posted on the main blog and not roto. I don’t play fantasy so I usually just skip the roto stuff due to a limited amount of time to read these articles, but occasionally they throw a gem like this on roto and I almost miss it. Anyways, awesome work Alex!

wildcard09
11 years ago

No worries, I’m sure for everyone like me who skips the fantasy stuff, there’s people who only read roto and not the main site. Glad I caught this piece though!

Chowjuch
11 years ago

Hey Alex, good work. I think your model is probably suffering from omitted variable bias however. Why only include Hard%? Certainly soft and medium contact effect BABIP too, just not to the effect of Hard (presumably). Additionally, if flyballs and infield flyballs have an effect so should groundballs. I understand the problem with including variables that add to one, but you can get past this by using ratios as variables. Try using GB/FB as a variable or soft contact/medium contact. And what is the point of Oppo% alone? Only to control for susceptibility to the shift? I think omitting up the middle or pull is a mistake. How about you create a variable that measures the total variation in a batter’s batted ball locations? It could be as simple as finding the standard deviation of their ball locations (a 200 to center, 200 to pull, 200 to oppo player would be have a standard deviation of 0 and would be the ultimate spray hitter). Likewise, why not use more than a year’s worth of data? It wouldn’t be difficult to run a pooled cross section regression on all the years we have Hard% for.

obsessivegiantscompulsive
11 years ago

Wow, Brandon Crawford has been hitting pretty well, but according to this, his BABIP should be much higher!

Kinda shocked to see Posey’s weak hitting, but that’s reflected in his drop in power so far this season, his poor hitting in cleanup has been a large part of the Giants offensive struggles.

Not too surprised to see Aoki and Panik doing well, right around what they have hit before for BABIP.

Justin
11 years ago

“I’m sure if we threw the kitchen sink at the problem, everything would fall into place. But the sink would probably break and it would be really messy and no one would want to clean it up and that’s why we can’t have nice things.”

Nailed it.

kyle Logan
11 years ago

can you use integrals and derivatives to find Bapip