Introducing the New New xK%, Featuring the Changeup Adjustment

Last week, I reintroduced my xK% equation, this time, with updated coefficients. The equation’s components were the exact same, so there was nothing new or exciting to report. However, on the following day, I published the top 10 over/ underperformers from 2011 to 2016 as a fun little exercise to learn who has broken the model. I then attempted to figure out if there was a common theme among the over/underperformers and after performing some additional research and calculations, settled on a possible explanation — changeups are bad for strikeouts. Turns out, I was actually onto something and that something was already discovered two years ago.

Frequent commenter Josiah Rutledge (whose commenter user name escapes me) emailed me to remind me of Eno’s article from January 2015 on this exact topic. Eno asked whether the changeup has a strikeout problem and the data he presents results in a resounding yes.

Since that validated what I thought might be missing from my xK%, I decided to go back to work. I added CH% to my data set and was encouraged to find that it’s correlation with K% was -0.184 during the 2011 to 2015 period. So far, so good. So I took the next step by rerunning the regression, and voilà, a new equation.

xK% = -0.8269 + (Str% * 0.2893) + (L/Str * 1.2515) + (S/Str * 1.1526) + (F/Str * 0.9454) + (CH% * -0.0282)

Adjusted R-Squared: 0.933

You Aren't a FanGraphs Member
It looks like you aren't yet a FanGraphs Member (or aren't logged in). We aren't mad, just disappointed.
We get it. You want to read this article. But before we let you get back to it, we'd like to point out a few of the good reasons why you should become a Member.
1. Ad Free viewing! We won't bug you with this ad, or any other.
2. Unlimited articles! Non-Members only get to read 10 free articles a month. Members never get cut off.
3. Dark mode and Classic mode!
4. Custom player page dashboards! Choose the player cards you want, in the order you want them.
5. One-click data exports! Export our projections and leaderboards for your personal projects.
6. Remove the photos on the home page! (Honestly, this doesn't sound so great to us, but some people wanted it, and we like to give our Members what they want.)
7. Even more Steamer projections! We have handedness, percentile, and context neutral projections available for Members only.
8. Get FanGraphs Walk-Off, a customized year end review! Find out exactly how you used FanGraphs this year, and how that compares to other Members. Don't be a victim of FOMO.
9. A weekly mailbag column, exclusively for Members.
10. Help support FanGraphs and our entire staff! Our Members provide us with critical resources to improve the site and deliver new features!
We hope you'll consider a Membership today, for yourself or as a gift! And we realize this has been an awfully long sales pitch, so we've also removed all the other ads in this article. We didn't want to overdo it.

Str% — Strike Percentage | Strikes / Total Pitches
L/Str — Looking Strike Percentage | Strikes Looking / Total Strikes
S/Str — Swinging Strike Percentage | Strikes Swinging Without Contact / Total Strikes
F/Str — Foul Ball Strike Percentage | Pitches Fouled Off / Total Strikes
CH% — Percentage of changeups thrown

If you recall (yeah, you definitely don’t), the R-squared from last week’s update was 0.931. So, this equation added all of .002 to the R-squared. It might seem like nothing and pointless, but a) hey, the R-squared did increase, albeit barely and b) the addition of CH% accomplished what I wanted it to, as you’ll see below.

Before getting to the individual players, a bit more results from the new equation compared to last week’s update:

New xK% Correlations
New xK% Last Week’s xK%
xK% Y1 to K% Y2 0.7173 0.7158
xK% 2015 to K% 2016 0.7137 0.7137
xK% Y1 to xK% Y2 0.7492 0.7475

So again, we see the tiniest of games from this new equation. It predicts K% in year 2 ever so slightly better (though somehow performed exactly the same from 2015 to 2016), while it’s marginally more stable from year-to-year.

Furthermore, CH% during the period was remarkably consistent, as its year-over-year correlation was 0.8772. It’s nice to add a component to the equation that doesn’t tend to bounce around each season.

Alright, now let’s get to the players and how, as a group, they have been affected.

New xK% vs Last Week’s xK% Group Averages
FBv CH% K% New xK% New K%-xK% Last Week’s xK% Last Week’s K%-xK%
Overperformer Averages (150) 92.3 7.50% 22.00% 20.53% 1.45% 20.46% 1.53%
Underperformer Averages (150) 91.2 10.80% 20.10% 21.89% -1.80% 21.92% -1.83%

So compared to last week’s equation, we see yet again that the new version offers the smallest of gains, but gains, nonetheless. Obviously, we’re not seeking a perfect match between K% and xK% for these two groups — there are always going to be outliers and every data point should never be expected to fit a regression model.

Now for the most important comparison, the changeup guys. On average, the pitchers in the entire data set threw the changeup 9.5% of the time. Below is a comparison of the 50 pitchers who throw changeups most frequently, along with all pitchers that have never thrown a changeup:

New xK% Changeup Group Averages
FBv CH% K% New xK% New K%-xK% Last Week’s xK% Last Week’s K%-xK%
Heavy Changeup Throwers (50) 90.7 30.2% 19.5% 19.9% -0.5% 20.5% -1.0%
Non-Changeup Throwers (103) 92.1 0.0% 21.4% 22.3% -0.8% 22.0% -0.5%

Woah, woah, woah, hold my horses. First, the good news — I set out to correct the underperformers after my theory seemed to indicate it was the changeup causing an issue. And that was essentially resolved, as the new xK% for the 50 heaviest changeup throwers is now much closer to the group’s actual K% as compared to last week’s equation. A reduction of 0.6% from the old to the new equation is pretty significant. Success!

HOWEVER, now the non-changeup throwers are all screwed up. They received an unearned xK% boost simply because they failed to throw even one changeup, and now xK% overrates the group even more than last week’s version (even when expanding the non-changeup throwing group to all very low users of the pitch still increased the gap between K% and xK%).

Check out the FBv (fastball velocity) column. My hope was that the non-changeup throwers would come in below the data set average (91.8 mph), which might suggest the need for incorporating that as a variable. But since the group sits above the average, all that would happen is my xK% would jump again and get even further away.

If nothing else, this table suggests that pitchers throw more changeups when their fastball ain’t so fast. That makes sense, of course. The correlation between CH% and FBv in my data set was -0.2126.

I’m not happy that while fixing the heavy changeup users, I seemingly broke everyone else. It’s kind of silly to feel the need to improve an equation that already sports an R-squared above .90, but hey, I’m a perfectionist.

Thoughts? Advice on how to fix the non-changeup users? Do you feel last week’s (without CH%) or this week’s (with CH%, and assuming no adjustment to fix what I broke) xK% should be the official version for my future posts?





Mike Podhorzer is the founder of ProjectingX IQ, an advanced fantasy baseball analytics platform that transforms projection data and in-season performance signals into actionable intelligence. He is the 2015 Fantasy Sports Writers Association Baseball Writer of the Year and three-time Tout Wars champion. He is the author of the eBook Projecting X 2.0: How to Forecast Baseball Player Performance, which teaches you how to project players yourself. Follow Mike on X@MikePodhorzer and contact him via email.

13 Comments
Oldest
Newest Most Voted
scotman144Member since 2016
9 years ago

Would it be against the spirit of your equation to put a conditional on things that could break it like throwing zero changeups? If you had (for example) an IF CH thrown > 100 threshold for that part of the equation to be included you’d solve the problem without really changing anything.

scotman144Member since 2016
9 years ago
Reply to  scotman144

More concisely: you could just wrap your old and new formulas in an IF statement in excel and hopefully get the best of both worlds.

=IF(CH% > [threshold value], new equation, old equation)

eldurkoMember since 2017
9 years ago

Maybe some sort of variable interaction. Just spitballing here, but CH%/FB% (which, I guess would just be a CH/FB ratio…). Or CH as a % of offspeed/breaking balls? Might help to change the scales.

snowybeard
9 years ago

Mike: is there any chance you could list the new tables for starters only? I play in two points leagues only and saves have some value but not as much as in cats leagues. (Also, these are keeper leagues, so the best relievers are kept.)
I plan to purchase your Pod’s Projections book, but only after I get out from under from some expensive dental and plumbing bills.

snowybeard
9 years ago
Reply to  Mike Podhorzer

Great! Thanks, looking forward to that.

I'm Your Huckleberry
9 years ago

I’m Josiah Rutledge.

Glad to know I’ve contributed to science in some small way.

I agree with the suggestion of setting some threshold of where the CH% factor kicks in. Throwing a change isn’t bad for strikeouts, relying on it is. You’d have to experiment and find out how many is too many when it comes to over/underperforming xK%.

You also could experiment with whether changeups per pitch or changeups per non-fastball is better

Saul Jackman
9 years ago

Really cool stuff! To generalize the conditional selection (IF changeup > threshold), you can interact the changeup term with each of the other four terms. That way you wouldn’t have to set an arbitrary threshold. Alternatively, if you want something even more general, you could use a non-parametric modeling approach (like a random forest) instead of linear regression. Partly depends on how many observations you have in the dataset though.

Saul Jackman
9 years ago
Reply to  Mike Podhorzer

Happy to instruct, if I can be helpful! Interacting the terms will be the more straight forward approach to implement. Effectively, you would add four coefficients (and input values) to your modeling specification. Your current model specification is this:

K% = b_0 + (b_1 * Str%) + (b_2* L/Str) + (b_3 * S/Str) + (b_4 * F/Str) + (b_5 * CH%)

The new specification would be this:

K% = b_0 + (b_1 * Str%) + (b_2* L/Str) + (b_3 * S/Str) + (b_4 * F/Str) + (b_5 * CH%) + (b_6 * CH% * Str%) + (b_7 * CH% * L/Str) + (b_8 * CH% * S/Str) + (b_9 * CH% * F/Str)

In words, what this does is combine your new and old model together. For those who throw 0% changeups, the model will drop out the last five terms, so it will behave exactly like the old version of the model. For those who throw changeups, it will be the new version of the model. There will be no threshold at which a pitcher is considered a ‘change-up pitcher’ or a ‘non-change-up pitcher’ — rather, the the more a pitcher throws change-ups, the more their strike rates will be modified.

To execute this, you just need to create four new input variables: (CH% * Str%), (CH% * L/Str), (CH% * S/Str) and (CH% * F/Str). Then add those variables to the model.

This method has its own assumptions, that you may or may not agree with. Mainly, it’ll assume that the effect of CH% on K% is linear — moving from 5% to 10% is treated the same as moving from 15% to 20%. Even if you don’t like this assumption, though, it’s at least worth trying.

If you are interested in this but not sure how to implement, I’m happy to talk more, just let me know if I can be helpful!

Saul Jackman
9 years ago
Reply to  Mike Podhorzer

Glad I could help! One thing to keep an eye on, though, is whether you’re overfitting the model. I’m not sure how many pitcher-years are in the analysis, but as you add more terms to the model, you run the risk of explaining the error term in the model. In practice, what will happen if/when you overfit is, your model will very accurately explain the K% of those you trained the model on (the in-sample group), but perform very poorly on out-of-sample predictions (i.e., predicting the future).

One way to check for this is to compare the adjusted R-squared of the model with and without the CH% variables — R-squared will always increase as you add variables, but adjusted R-squared may actually go down.