Archive for regression

The Biggest Hitter K% Outliers of 2018

Yesterday, I devised a new expected strikeout rate for pitchers and used it to identify qualified starting pitchers who over- or under-performed in 2018. I’m reluctant to make out the exercise to be more than it is. I simply wanted to take the most intuitive approach to describing a pitcher’s strikeout rate (K%): by using the plate discipline exhibited by opposing hitters. Today, I seek to do the same for hitters. I can tell you now the discussion will be much more qualitative than quantitative.

Read the rest of this entry »


2018 Statcast Park Impacts (Not Quite Factors)

The longer we have Statcast data at our disposal, the more ways we find novel uses for them. What follows is my proposal to use the difference between expected and actual value segmented by batted ball type and venue to determine park factors (and potentially evaluate defensive value, as described in the footnote). Unfortunately, someone smarter than me was already way ahead of me. I’ll get to that in a second.

A typical park factors grid, such as those produced by ESPN or FanGraphs, commonly relies on outcomes — outcomes of plate appearances (ESPN), batted ball categories, or both (FanGraphs). They describe what actually occurred, the way wOBA describes a hitter’s actual production. Conversely, expected wOBA (xwOBA) describes what should have occurred based on a batted ball’s exit velocity (EV) and launch angle (LA). It strips away everything else, holding constant all other environmental factors in order to deliver an otherwise-context-neutral EV/LA-based value.

The difference between wOBA and xwOBA (“wOBA minus xwOBA,” or wOBA—xwOBA for short), therefore, effectively captures all value amassed or lost by other variables. In other words, if wOBA explains what actually happened in a non-neutral environment, and xwOBA explains what should’ve happened in a neutral environment, then the difference between them characterizes the effect of the environment — the ballpark itself.

Unfortunately for me (but fortunately for everyone else), Tony Blengino already did this (which is why he’s a former MLB executive and I’m not). In 2017, he used Statcast data to calculate park factors on the basis of expected outcomes relative to actual league-average production. For all intents and purposes, it’s the same idea.

Consider this post a refresher on the topic.

Let me call your attention back to a simpler time. If you search “Miguel Cabrera xwOBA” on Twitter, you’ll find, well, not a multitude, but at least a sampling, of Tweets from the summer of 2017 lamenting Cabrera’s (and his teammate’s) bad luck by measure of wOBA—xwOBA:

Read the rest of this entry »


Predicted 2019 NFBC ADP

Disclaimer: This is just for fun. I am, by no means, claiming that the predicted average draft positions (ADPs) described below will happen. Obviously! I’m no prophet. Also, I am not claiming these predictions are merely educated guesses. In fact, these aren’t even my predictions — they’re yours. Or, well, they’re not your predictions — they’re my computer’s predictions, but fitting your behavior to observed events.

That’s a complicated way of saying: by using historical ADP data and end-of-season (EOS) values, we can model future ADP values. (xADP, if you will.) Namely, with 2018 EOS, 2018 ADP, and 2017 EOS, we can predict 2019 ADP — and explain almost 60 percent of its variance (adjusted r2 = 0.59).

Read the rest of this entry »


Madison Bumgarner’s Fastball is (Still) Broken

If something about Madison Bumgarner’s first eight starts of 2018 have seemed odd to you, it’s because they have been. No matter the fielding independent pitching statistic to which you subscribe — FIP, xFIP, SIERA (although, frankly, it should be SIERA) — Bumgarner’s 2018 has not inspired confidence. Despite a dazzling (and quintessentially Bumgarnerian) 2.90 ERA, his baserunner suppression skills (i.e. strikeouts and walks) have lagged this year, and the various FIPs all portend severe bumps in the road. Granted, Bumgarner has outperformed his FIPs the last three years and throughout his career. I’m here to argue not that we should dismiss our concerns because of this but, instead, that such overperformance has insulated us from what should be potentially serious concerns about MadBum’s long-term health and success.

The problems with Bumgarner’s 2018 season — or at least the peripherals that underpin his 2018 season — thus far stem back not to his broken finger but, rather, something both farther back and much more dire. You may or may not recall Bumgarner fell off a dirt bike last year and injured his throwing shoulder. He returned from that injury almost exactly a year ago and promptly underwhelmed us. Sure, he posted a 3.43 ERA through September and has a 3.23 ERA in the calendar year since his return. It’s not vintage Bumgarner, but it’s not awful. But the peripherals, oh, the peripherals: his strikeout rate (K%) has caved dramatically, falling more than 6 percentage points (27.1% from April 2015 through April 2017; 20.9% from July 2017 onward).

It’s his fastball. Bumgarner’s fastball, once elite (relative to other four-seamers), is broken, and it has been broken for a year.

Read the rest of this entry »


Diagnosing Jon Gray

In a fairly surprising turn of events, the Rockies demoted Jon Gray Saturday. Gray has arguably been baseball’s most enigmatic pitcher this year, posting a career-worst 5.77 ERA supported by career-best peripherals — e.g., a 13.4% swinging strike rate (SwStr%) underpinning a 28.9% strikeout rate (K%), and fielding independent metrics of 2.78 xFIP, 3.08 FIP, and 3.15 SIERA. Given our most basic sabermetric understandings of baseball, Gray should be a very good pitcher, even if he pitches half his starts at hitters’ paradise Coors Field.

I have written about how a common-breed Rockies pitcher’s peripherals might be penalized for calling Coors Field home (Gray inspired this bit of research as well). FIP metrics generally underestimate ERA by anywhere from 0.8 to 1.3 runs for home starts (compared to 0.0 to 0.2 runs for road starts), suggesting that Rockies pitchers may underperform (a) their FIPs by 0.35 runs or (b) their SIERAs by 0.65 runs — given error bars, maybe more.

Still, that doesn’t explain why Gray’s ERA is nearly 6 right now. I shed light on the ridiculousness of the move; his strand rate (LOB%) is suppressed and his batting average on balls in play (BABIP) is elevated, even compared to his uniquely bad baselines. I’m not sure there’s much more to it.

Nick Mariano of RotoBaller noted here that Gray’s fastball has been incredibly hittable since his debut and especially this year. Despite my thoughts on the inevitability of regression in Gray’s favor, I wanted to pursue Mariano’s train of thought a little further. Gray’s fastball is bad, but how bad? And why?

Read the rest of this entry »


Modeling SwStr% and GB% Using Velocity and Movement

This year, I’ve been caught up on pitching. I investigated the nuance inherent to swinging strikes, indirectly made a case for completely abandoning the sinker with this piece comparing pitch type outcomes, and (maybe) identified the keys to unlocking pitcher BABIP and HR/FB.

Here, I’ve modeled swinging strike and ground ball rates using only pitch velocity movement. Surely, this work can be improved; my quantitative tool set, while fairly robust compared to the layman, is meager compared to the professional or even hobbyist statistician. Regardless, I think it’s pretty cool, and I hope it adds to the conversation constructively.

Mostly, this serves to satiate my own curiosity. Unfortunately, it may be denser than I expected — few answers are ever quite as simple as you hope them to be, I guess.

Existing Research

I linked to several of my own pieces above. Dan Lependorf wrote about estimating ground ball rates in 2013 at the Hardball Times, although its conclusions have an anecdotal slant. (It thinks about velocity and movement but doesn’t take the requisite steps to bridge the logic.)
Read the rest of this entry »


ERA Minus SIERA Laggards: Gonzales, Archer, Gray

FanGraphs hosts a statistic for pitchers called ERA Minus FIP (“E-F”), which is as advertised. FIP being a (somewhat) adequate measure of pitcher over-/under-performance, one could look to E-F to identify pitchers who may, as they say, be due for regression. FIP’s correlation with ERA, however, is weaker than that of xFIP due to the former’s inability to account for the volatility inherent to home run-to-fly ball ratios (HR/FBs). To take it a step further, xFIP’s correlation with ERA is weaker than that of SIERA due to the former’s inability to account for a pitcher’s ground ball rate (GB%) and how it interacts with his strikeout and walk rates (K%, BB%).

Alas, I often use SIERA, rather than xFIP or FIP, to identify pitchers who may be ripe for regression. ERA Minus SIERA (“E-S,” henceforth) is not the be-all, end-all by any means, and I would never consider making a roster decision based exclusively on that metric. Player evaluation is a holistic endeavor, which you likely know yet I still intend to demonstrate. Three names stood out to me — four, if you include Luis Castillo, but I covered him a week and a half ago — as interesting E-S targets, but I came away from this feeling good about only one of them.

Read the rest of this entry »


Obligatory Monthly Update on Carlos Gonzalez

It’s time for another monthly analysis of Carlos Gonzalez’s tumultuous 2015 season. When we first tuned in, Carlos Gonzalez was bad. Like, really bad. Despite peripherals that suggested some bad luck, the rest detailed a hitter struggling mightily.

When we last tuned in a month later, Eno Sarris determined CarGo had been unlucky up to that point, but only slightly. Things had started to turn around, but it was hard to be optimistic.

We tune in now, another month later, to find Gonzalez falling short of his previous levels of production but still performing admirably considering the circumstances. To paint a fuller picture, observe CarGo’s statistics at the time we published each of the aforementioned posts:

Read the rest of this entry »


ERA-FIP, and the Importance of Situational Context

I like a lot of pitchers who have unperformed this year. With strikeout and walk rates (K%, BB%) of 20.5 percent and 7.0 percent, respectively, Drew Hutchison delivers everything I want from a mid-rotation fantasy starter. With a 5.19 ERA and a 1.47 WHIP, however, he delivers a flaming bag of feces to my doorstep.

The same can be said for Taijuan Walker who, after a terribly rough start to the season, dazzled for seven straight starts before recently tossing three stinkers. With plate discipline ratios better than Hutchison’s and just 22 years old, Walker demonstrates the skill set and ceiling that have earned him consensus top-20 honors on prospect lists from 2012 through 2014. Yet his 5.06 ERA and 1.29 WHIP have left fantasy owners not only disappointed but also reeling.

Hutchison and Walker share a common trait: their ERAs dwarf their fielding independent pitching (FIP) statistics. FIP was designed to demonstrate a pitcher’s true performance in light of the events he can control — that is, events independent of balls put into play at the mercy of the defense supporting him (among other things).

Read the rest of this entry »


An Expansion on xISO, Plus 10 Noteworthy Names

Last week, I introduced xISO, a metric that calculates a player’s expected isolated power based on his batted ball profile (per FanGraphs’ recently added batted ball data courtesy of Baseball Info Solutions). Having looked at a handful of underachieving National League outfielders for its induction, I’ll expand the analysis of xISO here today.

I’ll reiterate some key points. I used all 12 years’ worth of batted ball data for all player-seasons in which a hitter qualified for the batting title. The OLS regression specified pull rate (Pull%), hard-hit rate (Hard%) and fly ball rate (FB%) as explanatory variables and produced the following equation, which I deliberately omitted last week:

Read the rest of this entry »