Challenge #2: Prove that a Low BABIP = Inducing Weak Contact
Yesterday, I issued my first challenge. Sparked by Brandon McCarthy’s bizarre outing on Monday, I asked you to prove that his HR/FB rate was not bad luck. The challenge led to some great discussion, which is exactly what I had hoped it would do.
Now it’s time to move on to the second, and likely final, challenge. It’s a topic that I am more interested in and has been debated ad nauseam. Of course, I’m talking about pitcher BABIP. We have been taught that pitchers will tend to regress toward the league average, which has sat around .295 in recent years, as hitters actually possess the majority of control over how often balls in play falls for hits. So early on, we eventually came to accept this.
But similar to grasping at straws trying to explain a suppressed/inflated HR/FB rate, it has become fashionable to claim that a pitcher with a low BABIP attains such a mark because he induces weak contact, even when the analyst is fully aware of DIPS theory and agrees with it for the most part.
A low BABIP by itself cannot be evidence of the ability to induce weak contact.
There are two major problems with the assertion that a low BABIP means weak contact:
1) Weak contact isn’t necessarily a bad thing! Between bloopers, dribblers and good placement between fielders, there are ways to get hits without striking the ball hard. But of course, all else being equal, a pitcher would prefer a weakly hit ball to a hard hit ball.
2) A low BABIP could be posted for a myriad of reasons — a) balls could be hit right at fielders, making them easy to field, regardless of how hard the ball was hit, b) the defense could play its best due to pure randomness when the pitcher in question happens to be on the mound and c) yes, the pitcher could be inducing weak contact that makes it easier to convert those balls in play into outs.
The problem? We don’t know what combination of the above factors, or any others, is the true driving force behind the low BABIP. It could be just one, a combo of two, or some sort of mix of the three. We don’t know which factors are involved, nor how much credit each should receive.
But we do know some things. Like that the distribution of batted balls that a pitcher allows strongly influences his BABIP. The hierarchy of batted ball types in terms of limiting BABIP is thus:
IFFB > FB > GB > LD
So theoretically, a pitcher with an average defense behind him and a league average batted ball profile should produce a league average BABIP of around .295. And of course any deviations from the average should produce corresponding differences in expected BABIP. That’s really the only thing we know though. Trying to break down BABIP by pitch type is problematic because a pitch isn’t thrown in a vacuum, as it’s part of a series of pitches that all affect the end result of that pitch.
Intuitively, it might seem as if a pitcher who generates lots of swings outside the zone (O-Swing%), or allows lots of contact outside the zone (O-Contact%), which may very well be worse contact than balls struck inside the zone, would post a lower BABIP. Then even combining the two by multiplying them together — lots of outside the zone swings, and also lots of contact — would yield a meaningful correlation. But this is not the case. A quick trio of correlation calculations analyzing a population of 844 starting pitchers from 2005 to 2014 resulted in near-zero marks for all three metrics.
Some explanations offered for inducing weak contact include:
1) The quality of a pitcher’s stuff is so good that it induces weak contact. Batters simply have a difficult time squaring up the ball, because the pitcher is just so darn nasty. Yet, no evidence, statistics, data, proof is ever presented. It’s stated as fact, when instead it’s mere conjecture.
2) The pitcher excels at changing speeds which keeps hitters off balance and leads to weak contact. It sounds all well and good, but I have yet to encounter any actual research of this phenomena existing, with specific players examined.
Speaking of specific players, let’s pick one to dive into, because it’s easier to continue a discussion that way. From 2011 to 2014, Kyle Lohse has recorded a .269 BABIP (it was much higher prior to 2011), ninth lowest among 151 qualified starting pitchers during that time span. Lohse certainly isn’t a pitcher that springs to mind when you think of quality stuff. So that almost immediately eliminates the first possible explanation.
He has been a fastball (sinker)-slider-changeup guy throughout his career, but started mixing in his curve ball more often during the latter two years of the time period. Including the curve, his pitches ranged from the low 70s to around 90. He also threw his fastball only around 50% of the time, which is pretty low. His repertoire and frequency seem to at least make explanation two a possibility worth exploring.
His batted ball profile has been nearly a mirror image of the league average, with some minor differences. He has actually allowed a higher rate of line drives and a marginally lower rate of infield fly balls, but fewer grounders and slightly more flies. Overall, the distribution alone should lead to a BABIP around the league average. But it hasn’t.
And it surely wasn’t his defense either that should get any credit unless they happened to play significantly better behind him than for every other pitcher. From 2011 to 2012, the Cardinals ranked 26th in baseball in UZR. And from 2013 to 2014, the Brewers were almost exactly average, ranking 16th.
So I am issuing a second challenge:
Prove that Kyle Lohse’s BABIP is Due to Inducing Weak Contact or
Prove that a Low BABIP and Inducing Weak Contact Are The Same
Mike Podhorzer is the founder of ProjectingX IQ, an advanced fantasy baseball analytics platform that transforms projection data and in-season performance signals into actionable intelligence. He is the 2015 Fantasy Sports Writers Association Baseball Writer of the Year and three-time Tout Wars champion. He is the author of the eBook Projecting X 2.0: How to Forecast Baseball Player Performance, which teaches you how to project players yourself. Follow Mike on X@MikePodhorzer and contact him via email.
Right, this one is even harder… because even if you could show that a pitcher is getting weaker contact, you then have to prove he is actually INDUCING that weaker contact. This seems to fluctuate enough that you have to lean to the hitter on that one.
Insofar as IFFB are weak contact, I think that’s as far as you can go.
I think you can go a little farther.
If you think about the quality of contact, you can put it into these four buckets:
1. Swing and miss
2. Swing and almost miss (weak contact)
3. Swing and connect solidly
4. Swing and crush the ball.
With DIPS theory, #1 is reflected in the K% component and #4 is reflected in the HR% component, so pitchers with greater skill in inducing whiffs and avoiding meatballs will do better (allowing for ballpark effects), and this shows up in a better FIP.
BABIP, once you factor out the impact of defense and ballpark effects, would probably depend more on the hitter’s skill level than the pitcher’s, as you note. However, over the course of a season the quality of batters faced should become a less significant variable affecting the pitcher’s BABIP because they will have faced a broad cross-section of MLB hitters.
Now, if BABIP depends upon #2, #3, and the non-HR #4 buckets of quality of contact, it’s easy to see that there’s a possible luck component on OF flyballs that are almost HRs or barely HRs. That’s what led to the creation of xFIP. But for IFFB, the hitter has gotten so far under the ball that the swing is real close being in the #1 bucket, so those can be considered part of a pitcher’s skill.
And I would argue that there’s a similar argument that the weakest-hit grounders (topping the ball or hitting it inside the label or off the end) and also be considered near-misses like the IFFB, and would also be real close to being in that #1 bucket. I don’t have data like this available, but Tony Blengino has his secret stash of batted ball data and puts together columns about this all the time. The one he recently did about Dallas Keuchel shows that Keuchel somehow has both a ridiculous GB% and a below average quality of contact on ground balls.
So I think there’s a good chance that Lohse’s pitching style could lead to more weakly-hit fair balls than the average pitcher. It’s just hard to show or disprove without really granular batted-ball data.
See, I stopped reading when he suggested that there needs to be proof that “stuff” induces weak contact. I don’t disagree that if I wanted to “prove” it, I should produce some evidence, the issue is “what counts for evidence?” The bean counters believe you have to quantify everything, or it doesn’t count. I would counter that someone can quantify something, but have a measure with low validity (and never know it), and his “analysis” will be crap. OR, your analysis might be missing a variable or three that we might want to know about. Just because someone finds a correlation doesn’t mean we have accounted for all the variability. Personally, I have never doubted that BABIP is due in part to “luck” (where the ball is hit). The question is, have you explained everything?
IMO, this is getting awfully close to asking us to prove a negative. We know what we know about batted ball contact and that obviously plays a big part – esp. IFFBs. Without access to a large set of HITf/x data, I’m not sure we can do what you ask.
Brad, I think the point is to spitball, rather than prove. With math, statistics, correlations, whatever, you’re going to find more often than not that the “Why?” comes after the “Hrmm?”
We’re just as likely to get the answer to “Why does BABIP correlate to relative humidity” as we are to “Why does Kyle Lohse beat the BABIP Gods”
My favourite part of yesterday was reading some of the not so obvious theories.
Yeah… also what is the point of these articles unless they’re bringing some info to the table?
If a player hits into an out into a shifted defensive alignment, where before under a normal alignment it goes for a hit. has he made weaker contact? If said batter is constantly shifted against by opposing the defense and his babip goes down as a result, has that batters skills diminished?
The defense has improved, and the pitcher is rightly rewarded for the improved defense. This part is not defense independent
I couldn’t find anything earlier this year using “better” batted ball data:
In the “Batted Balls: Who Gets Credit Once the Ball is in Play?” section
http://www.hardballtimes.com/looking-at-pitcher-war/
I’m going to say Phil Niekro’s career offers some good evidence.
Now when you say Kyle Lohse doesn’t have the greatest stuff, how are you proving that statistically? When it’s put he’s not a guy who “springs to mind” when thinking of great stuff, from what numbers do you derive that opinion? Strikeouts. That seems to be the only “accepted” way.
Have we become (self included) so infatuated with strikeouts that we even need to ask if the ability to induce weak contact exists? Of course it does. That’s more or less asking if pitchers have the ability to deceive hitters either through pitch movement, sequencing or delivery. If a hitter is deceived, he is not likely to put a good swing on the ball and create anything but weak contact. Now would any baseball fan argue that pitchers can deceive hitters?
Now we know weak contact carries a lower BABIP than medium or hard contact across all batted ball types. So it would stand to reason that a pitcher would attempt to induce said contact, as even pitchers with super-elite K rates will still get more outs via batted ball than K.
I will say Im quite mathematically challenged so Id have no idea how one would prove such a thing with the data we have available. But we can’t assume that years of hitters have been wrong when they reference a pitcher being hard to pick up or hard to square up. There’s more to that puzzle than simply pitch movement, theres so many factors here it makes one’s head spin trying to think out where one would even start to look to prove/disprove such a thing with any reliability. Release points, deceptive deliveries, pitch sequencing and even intimidation could all be factors.
I personally believe pitchers have far more control over where the ball ends up than we have come to accept. Not all pitchers possess said skill, and I would tend to think it’s an “old players” skill. I think it takes a high pitching IQ and not only good stuff, but the right mix of stuff to be a pitcher who can in fact do it consistently. Logically it would make little sense that the batter has more control than the pitcher, at least to me, because hitting is reactionary, the pitcher initiates the encounter. He decides first where he’s throwing the ball, how, and does so the batter is only left to guess and/or react.
We often look at the pitcher as the defense and the hitter as the offense, and in the game sense they are, but in reality their roles during the batter/pitcher matchup are reversed, the pitcher is on the offensive. So it would stand to reason that they control the result at least as much as the hitter does.
Sorry for the length/hard to read.
No need to apologize fathead, very well stated.
I love the reminder about the pitcher being on the offensive in the batter/pitcher matchup. Hitters vary in their ability to square up on the ball and in the distance they can make the ball go, but a pitcher with great stuff, command, or deception can fool hitters of all types often enough to blunt their effectiveness, whether by a swinging strike or a ball not hit squarely.
Here’s an article from Blengino about this, although he’s focusing on Dallas Keuchel:
http://www.fangraphs.com/blogs/dallas-keuchel-and-weak-contact/
It doesn’t necessarily explain all the reasoning, but he talks about how the biggest thing for Keuchel is being the league’s top grounder-inducing pitcher at 60%, and allowing nearly the lowest liner totals. He says the way Keuchel achieves this is setting up his sinker that has the same velocity as his 4-seamer with late sink that is the key to batters not being able to lift the pitch.
A pitcher’s control can make up for lack of stuff. Greg Maddux, had a BABIP of about .275 during his peak years in the steroid era. I would think that normally, better control will lead to lower BABIP. Also, low BABIP could be caused by avoiding hard contact just as much as it could be from inducing weak contact.
Johan sanatana, career .276 babip, 13% IFFB%, 9% hr/fb%. He’s (or was) a pretty good weak contact guy.
Yeah but what about Johan Santa?
The extra A is for a lower Average on balls in play.
lol
As a side observation, if good stuff is the cause of bad contact, is there a correlation between k% and low BABIP?
Take a look at the stats of the very best pitchers, year in and year out..don’t neglect to notice the home runs allowed either..
And do what, exactly? Take a look at the worst pitcher’s stats. They should be just as informative – maybe even moreso.
1st thing we have as evidence of weak contact is his HR/FB rate. He allows a lot of flyballs, but comparatively few of them have gone over the fence. We know he pitches in Milwaukee for a chunk of the time and that has a reputation for increased HRs according to park factors.
2nd thing we have vs average MLB is that Lohse, despite giving up a ton of Flyballs, actually pitches lower in the zone than the composite MLB pitcher. To me, this says hitters are “trying to lift” the ball, or his sinking action encourages that they swing slightly under it. Why else does he give up so many flyballs? I’d bet he has 2 different Sinkers… or conceptually more sinker/2seamer where he is getting weak contact because these 2 pitches are so similar but his 2 seamer tails more than sinks.
3rd thing is again… heatmaps. Lohse is consistently down and in on almost everybody. Average gets more of the plate. Overlay the ISO/P heatmaps against Lohse’s / MLB average. The zones he is most consistently in correspond to weaker (measured by ISO) than average contact.
frig you lahey
Jim Lahey is a Drunk Bastard!!!
If you take the overall results of a group of pitchers, it will correlate that the guys who keep the ball in the park and give up fewer line drives vs. the rest of the group – those guys will have the lower pitcher BAPIPs.
So, yes also saying pitcher’s HR% or HR/FB% is a bellweather for other in the park power (line drives).
—
Fancy baseball stats are usually not mathematical proofs.
Can you offer some proof of this?
Just some ideas here, that I don’t know (or have the skill) to attempt to prove:
1) Using standard deviations, can’t we figure out that a certain number or percentage of pitchers are likely to have an abnormally low (at least 1 SD) BABIP for 1, 2, 3, even 4 years in a row, even just if BABIP is absolutely all luck-based? We have no way to predict which pitcher it will be, but it’s possible someone like Matt Cain or Kyle Lohse are just outliers because, statistically speaking, there has to be an outlier. It just happens to be them.
2) How do we account for shifting? If every team shifted in an optimal way at all times, BABIP would probably drop league-wide. But since no team shifts in an optimal way at all times, and some teams barely shift at all, it’s possible that some teams are being inordinately helped by the shift.
3) What about pitchers who throw a fastball/sinker/cutter that tails away from the plate? If a righthanded pitcher throws a fastball/sinker/whatever that has glove side run, it will obviously tail away from righties. If this pitcher tends to throw on (and just off) the outer half of the plate a lot, that pitch may be really hard for righties to square up correctly. Obviously, that would be the opposite of lefties, and presumably righties could eventually wait on that pitch and still hit it. But I wonder if a pitch tailing off and away is harder to get good contact (and presumably more likely to have a lower BABIP)?
4) Does a large differential between a fastball and breaking ball and/or changeup make a difference? Do guys with 10 mph between fastball and breaking ball, and another 10 between breaking ball and changeup, tend to have lower BABIP? Do guys with a smaller difference have a higher BABIP?
5) Does pitcher height matter at all? What about sidearmers or three-quarters guys? What about guys like Carter Capps and his funky motion?
6) Some people above mentioned great control possibly correlating with low BABIP. Can we check this with pitchfx data or even just BB rates? If a pitcher (like Maddux, who was mentioned above) can consistently hit just off the outside corner, maybe that leads to a lower BABIP?
7) Are there guys who have thrown a lot of innings and have a really consistent split in BABIP rates vs. righties instead of lefties? Does that tell us something about their delivery or the run/cut of their pitches being more effective against one side or the other?
I know that none of these are answers, just trying to think out loud for now.
2) Yes, some teams are doing the defensive infield shifting better than others and getting more outs because of it. It definitely helps.
In theory, all you’d need to do is determine defense for each pitcher. That would give you an idea of what BABIP you’d expect each pitcher to allow (those with good defenses would allow a lower BABIP and those with bad defenses would allow a high BABIP). Then you’d be able to determine a sample of pitchers that consistently overperform and underperform their expected BABIP over a few seasons and that would give you something to test.
It just requires defense data for each games which isn’t available to the public. But someone that has that could figure it out. Lacking that, I like my theory.
Greg Maddux had about a career .275 BABIP during his peak years. He was never a strikeout pitcher but he is among the greatest pitchers of all time. He used smarts and control. Therefore intelligence and control can reduce BABIP.
With Maddux and Loshe(how the 2 got in the same sentence is crazy), the low BABIP has to be partiality from inducing weak contact and/or avoiding hard contact.
Low BABIP and weak contact are not the same simply because a lower hard hit contact rate will have a similar effect as a higher weak contact rate. Same would be true if a pitcher had both a higher hard contact and weak contact rate than average.
Without doing days and days of research I don’t think you can get the whole story behind Loshe’s low BABIP. For one you would need to go more in depth to show how much better control and intelligence affect contact rates. There are several others factors to consider. For example, as a sinkerballer will induce more groundballs, what is to say that Loshe’s pitches don’t cause more batters to hit the ball to a certain place? The predictability of where the ball will be hit can bring down BABIP without a change of contact rates.
Yes, but Greg Maddux was a Hall of Famer for a reason: he had other worldly control. As Jeff Sullivan put it into a piece about him last year after he was inducted into the HOF: “…somebody once said about Greg Maddux is he could throw a ball into a teacup.”
Maddux had 4 seasons during the steroid era where he gave up fewer than 10 HRs, including one season where he only gave up four in 202 IP and another where he gave up 7 in 368 IP. That indicates weak contact indeed.
Greg Maddux also had the benefit of good defensive teams with the Braves, something out of his control.
I just don’t buy the intelligence thing as being a huge factor in inducing weak contact. Mike Mussina was one of the brightest pitchers to ever take the mound, yet his BABIP fluctuated greatly between 1995 and 1997 (.251, .318, .282, respectively). And, as evidenced by his outstanding strikout to walk ratios, he had very good command of his pitches too. But the Orioles of the 1990s didn’t have great defensive teams behind him. It also didn’t help that he had to face the DH and pitch in Camden Yards.
I’ve often wondered what Mussina’s numbers would’ve looked like had he pitched in old Memorial Stadium and had Paul Blair in CF and Mark Belanger and Brooks Robinson playing behind him. I think it’s safe to say that his BABIP would have been consistently lower.
Maddux is the extreme example, but the point is he did not have the greatest stuff but still posted a lower than average BABIP. Granted he did have a great defense, but that cannot be the only reason.
BABIP fluctuation will happen. Maddux posted a .324 one year. Doesn’t that fact that he posted slightly below average BABIP with a bad defense indicate that he has a skill?
The basic issue is that there is a ton of research that needs to be done. So many things need to be taken into account: defense, opponents faced, control, intelligence, stuff, etc.
I don’t think that’s correct. I believe that major-league caliber pitchers who give up weak contact and strong contact have roughly the same BABIP (those that give up weak contact have a slightly lower one).
In the below article, I note the performance of batters based on the pitch count. Basically, pitchers have similar BABIPs regardless of whether they give up contact on a hitter or pitchers count but they give up harder contact (as measured by extra base hits) in hitters counts. I propose this means that pitchers that give up weaker contact (like Kimbrel or Rivera) allow fewer extra base hits but more singles and therefore we don’t see the impacts in BABIP.
camdendepot.blogspot.com/2015/03/fip-and-ball-strike-count.html
It’s incomplete to talk about extra base hits in context with BAPIP.
Home Runs are thrown out of the sample, and these are extra base hits, too.
For BAPIP a line drive base hit rip does count, line drive home run, no.
That’s the point I’m making.
Pitchers that give up weaker contact (allow contact in pitchers counts) and those that give up harder contact (allow contact in hitters counts) have a similar BABIP because the ones that give up weaker contact allow fewer home runs (not included in BABIP), fewer triples (included in BABIP), fewer doubles (included in BABIP) and more singles (included in BABIP).
Basically, the increase in singles allowed makes up for the decrease in doubles and triples allowed and therefore results in a minimal difference in BABIP.
Therefore, if I’m correct then weaker contact won’t result in a lower BABIP.
No you are still grouping the Ichiros with the Pujols..it’s not fair.
Pujols hits 35 line drive homers that Ichiro hits for singles. One is counted in BAPIP, another is not.
You are just working around the BAPIP stat parameters, where home runs are not there..
I don’t understand your argument. We’re talking about pitchers and not hitters.
The challenge asked whether a low BABIP was related to weak contact. I argued based on results via pitch count that it wasn’t.
You stated that low BAPIP (Bases Allowed Per Innings Pitched) is related to weak contact. I agree that you are correct (I misunderstood you earlier) but I’m not sure why your statement is relevant to the challenge. He asked about BABIP and not BAPIP.
The point is that everyone uses BABIP and not BAPIP. So, that’s why ignoring home runs is useful. Because one would think that a pitcher that gives up weaker contact would allow fewer hits on balls in play. If you want to argue that BAPIP is related to weak contact then that’s absolutely the case. But the reason why FIP includes home runs is because everyone agrees that pitchers can control home runs. And a pitcher that gives up 50 home runs is dinged more than a pitcher that gives up 5 (presuming that they pitch an equal amount of innings).
So why does it matter that I’m ignoring BAPIP?
I won’t say that I can prove weak contact for Lohse, but I think we can see some evidence to support it. I would look at SLG-against on all fair balls as a proxy for hard contact. Here is Lohse’s SLG-against on fair balls compared to the NL average SLG-against:
2014 .495 NL avg.: 509
2013 .497 NL avg.:.511
2012 .465 NL avg.:.521
I think if you look at similar comparisons for pitchers believed to be elite in contact management, you will get somewhat larger margins relative to the league average SLG on fair balls. So, maybe he isn’t in the elite category for contact management. But, over the last three years, he has consistently recorded a lower SLG on fair balls than the average pitcher. (Note: I used B-Ref splits pages.)
Because Lohse did not have a high GB percent over that time period, the GB ratio doesn’t bias his SLG-against downward. If anything, his Flyball ratio of 36%-41% over that period should lead to an above average SLG-against ratio.
I won’t say that I’m convinced that Lohse’s contact management skills have suppressed his BABIP, but there is enough there to warrant further investigation.
Dare I say, why not take a look at ISO against to get an idea of the type of contact pitchers are allowing?
As I recall, SLG stabilizes more quickly (smaller sample) than ISO, and also correlates with next year SIERA and FIP more than ISO.
Baseball Info Solutions seems to be doing some very innovate work around Trajectory-Based Hitting and Pitching statistics. I think some insight into the answer to this challenge lies there.
Just taking a cursory look at the statistics, i think it’s useful to introduce survivor bias on purpose. Year-Over-Year on BABIP is going to be squat, and even with some rather silly fudging of the numbers, we’re not going to be able to get anything better than skill – maybe – accounting for 10% of the correlation between anything and BABIP. So when you’re looking at a *good* pitcher being maybe, if we’re lucky, 10% better than the average and you’re saying that, various qualities surrounding one’s change-up influence BABIP to an absolute maximum of 5%, you’re basically left with… well, nothing.
There are a few correlations worth noting, and off the top of my head, having a change-up with a decent speed difference and a couple pitches moving on the x plane, seem to come to mind.
Almost every study on BABIP has been done, rightfully so, on all pitchers which is going to nullify a lot of the stuff we’re looking for — If we take the entire sample set, we’re basically weighting everything equally. As I said when I started this off, purposely introducing survivor bias is going to allow us to see at least some correlation year-over-year. Whether you’re setting the minimum to 100 or 150 or whatever, you’ll get at least a bit *more* correlation.
Needless to say, when you’re getting the numbers you get for batted ball profile year over year and then BABIP year over year, you can kinda see why people just regress BABIP to the league norm.
Personally, I like to believe that at least a couple pitchers have figured out the secret sauce for BABIP. But there’s no real way to separate them from the noise.
I was looking at some numbers for this and while I am not sure it is relevant to weak contact, I was wondering if someone with better leaderboard sorting ability/mathematical skills/something could help me with this, though it might be nothing.
I was looking at First strike %/zone %, because I knew Lohse had a high first strike % and I was thinking that it was possible that was related to his low BABIP, the logic essentially being that the more you get ahead, the more likely you are to have people “defend” the plate and such which leads to worse (not necessarily weak, just worse) contact or what have you. I added in zone % because I thought the rate that one throws balls in or out of the zone may affect this.
My first intuition was high fstrike% + zone % would lead to some lower BABIPs, but when I looked at it, a different pattern seemed to emerge, which is that a high fstrike% with a lower zone %(50% or so) seemed to lead to lower BABIPs. From 2011 to 2014, the same years of Lohse mentioned here for consistancy, I looked at players in the Top 60 in frstrike% (61.1% or higher) and saw which ones had a zone % roughly equal or less than Lohse’s than Kyle Lohse’s and their BABIPs, this is what I came up with. FStrike% listed first, Zone % second, BABIP third:
Kyle Lohse: 66.7%, 50.1%, .269
Tommy Milone: 66.8%, 46.3%, .296
Patric Corbin: 65.7%, 50%, .295
Josh Tomlin: 65.6%, 51.3%, .287
Clayton Kershaw: 65.5%, 51.3%, .264
Kris Medlen: 65.1%, 50.9%, .283
Jose Quintana: 64.8%, 49.3%, .300
Jake Peavy: 64.2%, 50%, .284
Chris Capuano: 64.1%, 49.3%, .306
Dan Haren: 64.1%, 47.8%, .286
Homer Bailey: 64%, 48.1%, .289
CC Sabathia: 63.9%, 49.1%, .309
Roy Halladay: 63.9%, 49.6%, .293
Wandy Rodriguez: 63.8%, 49.3%, .282
Hisashi Iwakuma: 63.8%, 49.9%, .271
Matt Garza: 63.7%, 48.9%, .286
Ervin Santana: 63.2%, 48.8%, .275
Adam Wainwright: 63.1%, 49.6%, .295 (Top 30 FStrike% ends here)
Felix Hernandez: 63.1%, 48.1%, .296
Joe Blanton: 63.1%, 50.2%, .330
Gavin Floyd: 63%, 49.4%, .292
Anibal Sanchez: 63%, 50.%, .303
Stephen Strasburg: 62.9%, 48.5%, .295
Kyle Kendrick: 62.7%, 48.5%, .287
Tim Hudson: 62.6%, 47.3%, .281
Dallas Keuchel: 62.5%, 45.1%, .307
Rick Porcello: 62.4%, 49%, .318
Justin Verlander: 62.4%, 50.9%, .286
Freddy Garcia: 62.3%, 46.8%, .286
Matt Cain: 62.2%, 50%, .260
Eric Stults: 62.2%, 45.6%, .290
Jason Vargas: 62.1%, 49.3%, .285
Mike Minor: 61.8%, 50.4%, .290
Mat Latos: 61.7%, 49%, .281
John Danks: 61.6%, 51.3%, .295
Of the 35 pitchers I listed (I dunno if top 60 is too high a sample sized, hence why I mentioned a cutoff at top 30), 23 of them had a BABIP below the league average of .295, 13 of which were also below .285. This by itself might not mean anything, but I did find it interesting, so I had some questions for those more statistically inclined/leaderboard inclined than me:
1. If I wanted to run some mathematical on how it might correlate to BABIP, like R-Squared or whatnot, how would I go about doing it? Is this something you can do that with?
2. This test didn’t control for pitchers with similiar makeup: Stephen Strasburg obviously pitches his % differently than Kyle Kendrick, even if their %s here are close. Further study may wish to do this.
3. Probably beyond doing effectively, but I feel a large part of this may be sequencing. You get your first strike in, then you can do stuff just outside the “zone” while ahead (and thus have a lower zone %) and cause people to hit the ball with poor timing, poor placement on hitting the ball etc, thus inducing “worse” contact. It is also possible that the lower BABIP is a skill, but is not necessarily inducing weak contact, but instead doing it in other ways (For example, IFFBs hit hard, or perhaps in ways that make them more likely to be hit towards defenders, thus making them more likely to be caught even if they are hard hit?).
4. Is there any way to search for high FStrike%s + Zone %s under a certain amount on the leaderboards? I’d be interested in this.
5. In general this is something small but I am interested in it, so general feedback is appreciated, even if I am sure I did silly stuff here.
These challenges are fun, even if I feel I’m underqualified to try and answer them. I will say that, with both this and the first challenge being for pitchers, I would love to actually see a Challenge #3 but about hitters or a hitter or whatnot.
Thanks for looking into this. I quickly ran correlations and a regression using 754 qualified starter seasons from 2006-2014. Correlation of Zone% with BABIP and F-Strike% with BABIP. Both were tiny, Zone% at .06 and F-Strike% at .03. It’s both a positive correlation meaning more strikes equals higher BABIP, but obviously the effect is minuscule.
Combining the two, R-squared is just .005, so nothing to see here.
Well, thanks for running it anyway, even if nothing really notable popped up. It seemed like a fun idea to check out at least.
The easiest way to do stuff like this is to head over to Fangraphs Leaders and create a Custom Report with everything you’re after (ie. just use everything with whatever limits you want)
Import it into excel or open office or whatever, and do some formatting so that the percentages actually are usable as percentages and not text strings.
After that, you’re free to use CORREL() or PEARSON() or COVAR() or whatever you darn well please.
There’s also six years of data on a fangraphs page, I always just google “Fangraphs all the pitcher correlations” and that’ll bring you to the tool. I think it’s from 2013.
There is also a bunch of free software out there for stuff that involves dicking around. I use R. You just import the CSV, normally everything goes smoothly. I’ve always used Linux, but I can’t imagine the package wouldn’t be widely available for windows. The default settings for almost every chart/graph/statistical model will work fine for you. You don’t need a tonne of statistical knowledge to get shit poppin’
Thanks for the help, Kristopher. 🙂 I’ll try and find something like you said next time I get some data I’m very interested in.
How can you possibly expect to prove that a pitcher induces weak contact without hit/fx data? The closest one can come right now is FB%, GB%, LD%, IFFB%. The data available to the public now cannot be used to definitively prove that some pitchers run low BABIPs because of weak contact, because we can’t even sufficiently measure the quality of said contact.
EXACTLY! That’s the whole point of this challenge. I’m tired of reading analysts claim that Pitcher X has a low BABIP, but it’s sustainable because he induces weak contact. It’s stated as fact, even though we have no way of proving it right now.
It would be interesting to look at BABIP for different counts as I would expect it to be lower with two strikes and when the pitcher is ahead in the count.
Here is a high level summary of BABIP by count for 2011-2014:
Who’s Ahead MLB Lohse
Pitcher .289 .273
Even .298 .275
Hitter .303 .268
So there is a moderate improvement in BABIP by count for the average MLB pitcher. However, Lohse doesn’t show any real variation, in fact he’s allowed a slightly lower BABIP on balls put into play when he’s behind the hitter (over 935 PA).
I have hunch on Lohse’s low BABIP when behind in the count. Let’s compare what happens to Lohse on a hitter’s count vs. Craig Kimbrel and Gary Gopher:
Kimbrel: throws filthy pitch, hitter swings and misses.
Impact = none to BABIP, but moves closer to possible K.
Gopher: throws meatball, hitter squares up on it and hits it hard.
Impact = even after factoring out the higher HR% in this situation, likely to produce BABIP above the league average
Lohse: throws a strike that looks good to hitter, hitter not fooled enough to whiff on it but IS fooled enough to hit the ball in play weakly.
Impact = if for every 25-30 balls in play you do this just once more than the average guy, you get a BABIP about .035 better than average.
The more I think about it, the more it makes sense that there is a class of guys with decent stuff plus above-average command who excel at fooling hitters, but who get weak contact instead of swinging strikes because of their lack of dominant stuff. I’ll explore this further, but I wouldn’t be surprised if Lohse fits that mold.
Another thought I had, but I couldn’t find the data to back it up…I recall a while back in some FanGraphs article, I couldn’t remember which one/find it, talking about BABIP on each pitch, and if I recall correctly fastballs had the highest BABIP of all pitches. The only other article I could find relating to this is a very old and thus likely not relevant Hardball Times article (http://www.hardballtimes.com/fastball-slider-changeup-curveball-an-analysis/).
If fastballs have higher BABIP than other pitches then, while it obviously would not account for most of why his BABIP is .269, it could be a small part of the puzzle of why Lohse has a lower BABIP than usual (For example it may mean his “expected” BABIP is .290), simply because he is using a pitch mix with a lower expected BABIP (His FB% is 45th lowest in the league from 2011 to 2014 among qualified).