Toward a Pitch Arsenal Score Statistic
You’ve heard me yammering about pitch-type peripherals for two years now, and we’ve made some advancements along the way. We established some good pitch-type peripheral benchmarks, and we took a first look at properly weighting each pitch. We’ve started to get a sense of how these things interact when it comes to the shape and speed of pitches. We’re making progress.
It’s worth stepping back and figuring out what the aim is at this point. Because we aren’t trying to rank the best starting pitchers overall, really. We’re trying to find undervalued pitchers before the market realizes that they’re good. So we have to move in the smallest possible samples. And we want to have a list of great pitchers that has some weird names on it as well. Those names, we hope, will soon start to make sense.
So, to that end, I’ve taken each pitch type and looked at only those pitchers that have thrown 100+ in each of those types. I’ve summed the ground-ball and swinging strike rates for each pitch, and then found the standard deviations. I’ve given each pitcher a z-score for his ground-ball rate and swinging strike rate on each pitch type. Then I’ve summed the z-scores for each pitch type, and then for each pitcher.
What we should be looking at is an Arsenal Score. With this way of looking at things, it’s possible to have one dominating pitch and still score well. Or a group of lesser pitches that are all positive.
What we haven’t done yet is nail down what the smallest sample for each pitch is. Or how to weight the pitches. Or how to weight the whiffs versus the grounders. So this may look different once we weight each pitch differently, and if we find a way to weight grounders and whiffs more correctly.
But at least we have a first attempt at it here.
So here we go. Here are the top 30 pitchers by Arsenal Score, or the sum of their pitch-type z-scores for grounders and whiffs.
| Pitcher | Total |
|---|---|
| Carlos Carrasco | 8.93 |
| Carlos Martinez | 6.81 |
| Felix Hernandez | 6.78 |
| Gavin Floyd | 6.77 |
| Marcus Stroman | 6.12 |
| T.J. House | 5.58 |
| Jaime Garcia | 5.54 |
| Brett Anderson | 5.24 |
| Zack Wheeler | 4.93 |
| Corey Kluber | 4.88 |
| Clayton Kershaw | 4.69 |
| Francisco Liriano | 4.07 |
| Cole Hamels | 3.74 |
| Trevor Cahill | 3.55 |
| Brandon McCarthy | 3.53 |
| Homer Bailey | 3.51 |
| Gerrit Cole | 3.51 |
| David Hale | 3.43 |
| Garrett Richards | 3.40 |
| Jacob deGrom | 3.39 |
| Tyson Ross | 3.23 |
| Wade Miley | 3.19 |
| Allen Webster | 2.76 |
| Gio Gonzalez | 2.75 |
| Sonny Gray | 2.71 |
| Jake Arrieta | 2.58 |
| Yordano Ventura | 2.48 |
| Wily Peralta | 2.39 |
| Tyler Skaggs | 2.37 |
| Jeff Samardzija | 2.36 |
This is limited to starters — Bryan Morris was actually at the very top of this list before we culled the relievers. And that might be a bit of a problem. If you only threw two pitches, but both were really good, you’d do well here. Should you? Or should you get a demerit for not throwing more pitches?
Carlos Carrasco doesn’t care. Either way, he was there in the top five, the lone starter among relievers. Each of his pitches has rock star peripherals, and he blows the field away. But Carlos Martinez was a reliever last year, and maybe he shouldn’t even be here. He didn’t throw 100 changeups, so we’re judging him solely on the basis of his breaking ball and two fastballs (the sinker was second-best to Aaron Sanchez with a 3.2 score, minimum 240 thrown, and the curve was third-best overall to Brett Cecil and David Robertson with a 2.5 score, minimum 350 thrown). Which are awesome. Does it matter that his change that has diverging peripherals (17.5% swinging strikes, 42.5% ground balls, 40 thrown)?
Felix Hernandez is great, so nice that he’s near the top. Ditto Clayton Kershaw, Cole Hamels, Gerrit Cole, Jacob deGrom, Gio Gonzalez, Sonny Gray, Yordano Ventura, and Jeff Samardzija. Call that the Affirmation Wing of the list. You knew they were good.
The veteran question marks often have the Carlos Martinez question mark behind them. Gavin Floyd‘s changeup doesn’t show here, even if he found something with it this year. Brett Anderson is graded on his sinker/slurve alone. Ditto Brandon McCarthy’s lack of change. Tyson Ross samesies. Even Sonny Gray has the changeup problem.
But there are some names that should go on your sleeper lists because they are listed here. Marcus Stroman has the arsenal of an ace. T.J. House got legit results on his many pitches, and he survived a negative sinker (-.1) this year and thrived. Zack Wheeler also did well despite a changeup that rated negatively by this metric (-0.6). Brandon McCarthy had the best four-seamer by a starter (2.3), and adding that to his two-seamer/cutter/curve mix might be all he needs for an excellent year in 2015. Wade Miley might be the odd pitcher to do better in the American League — all four of his pitches rated well, and his slider (1.7) was a top-fifteen slider among starters.
Then you get to the guys that should be deep league considerations near the end of the list. David Hale will get even more credit if we find a way to reward balanced arsenals better — he has four pitches, and three of them had a combined z-score of one or better. If the Braves get another starting pitcher, though, he won’t have a starting job. If he goes into camp as the fifth guy, though…
Allen Webster has terrible command. And terrible body language. These are things that were said about Carlos Carrasco, but again, they’re there if you watch him pitch. And yet his zone% is only ten percent worse than league average, and his first strike rate is virtually league average, and he had modest walk rates in the minor leagues. And his arsenal includes the fourth-best changeup by z-scores (2.4), an above-average slider (.15), a good sinker (2.0) — and survives a terrible (-1.8) four-seamer. What if he ditched the four-seamer and only walked four batters per nine innings next season? He’d still have to leap over a lot of names on that depth chart.
So there you have it. A first shot across the bow. A name. A first round of sleepers.
Perhaps soon, Arsenal Score will get an update.
With a phone full of pictures of pitchers' fingers, strange beers, and his two toddler sons, Eno Sarris can be found at the ballpark or a brewery most days. Read him here, writing about the A's or Giants at The Athletic, or about beer at October. Follow him on Twitter @enosarris if you can handle the sandwiches and inanity.
So what about weighting the pitches on how often they are thrown?
For example, if a pitcher has an above average fastball that is thrown 60% of the time, it should matter more than a breaking ball that is way above average, but only thrown 10% of the time.
This is a good idea.
And what about finding the mean value of a pitcher’s pitch instead of adding it and giving so much credit for a pitcher with 4 slightly above average pitches that he only uses two of? 😀
I think there is some value for weighting and non-weighting.
Weighting will get closer to xFIP (strikeouts and groundballs).
Non-weighted looks to find pitchers who could possibly change their pitch mix to see some improvement.
This edges even closer to tripping over the elephant in the room, which is sequencing. One can imagine two pitchers with exactly the same repertoire thrown exactly as often, but with very different results because of the order in which they throw the pitches.
Wouldn’t the sequencing factor already show up in the whiff and GB numbers?
Agreed. This would make sense as a counting stat; sum of the z-scores of all pitches thrown. (Or as a rate stat weighted by percentages I guess.) As it exists, I think this system probably overvalues guys who have a lot of pitches and thus more z-scores to sum.
It seems like pitch weighting should be done on the basis of how important of a pitch it is in the game of baseball (which is somewhat dependent on how often the guy throws it, but not totally)
Having a bad fastball is, in my mind, harder to overcome than anything else. So it should be weighted accordingly.
Also, a lousy curve hurts you less if you have a filthy slider. But I would think having both is probably more of a boon than just summing the curve’s value plus the slider’s value.
So I did this same thing a while ago, except I incorporated velocity, movement (for some pitches), BB%, and I weighted the z-scores by how often the pitcher throws that pitch. Here’s my top 10 starters:
Jose Fernandez 4.22
Clayton Kershaw 3.37
Corey Kluber 2.94
Masahiro Tanaka 2.92
Yusmeiro Petit 2.66
Felix Hernandez 2.66
Derek Holland 2.52
Marcus Stroman 2.51
Chris Sale 2.51
Stephen Strasburg 2.39
and top 10 relievers:
Mark Melancon 8.26
Kenley Jansen 6.85
Zach Britton 6.46
Tony Watson 6.37
Jason Lane 6.20
Brett Cecil 5.68
Chad Qualls 5.57
Eric Jokisch 5.49
Carter Capps 5.30
Sergio Romo 5.26
Derek Holland is shocking. It is amazing that his Slider saved him THAT much.
It’s mostly because of his 71% GB% on it, not sure I would expect that to continue.
this is a much better list, the one in the post is absurd
Honestly, I’m looking for players that look absurd. Because it was this analysis that netted me Carrasco, Hahn, Skaggs, Samardzija, Keuchel, and more sleepers last year. If it’s all guys you like, it’s not useful.
Also, I’m not sure about adding in velocity: there’s a ton of multi-collinearity issues there. Velocity affects swSTR, for one. I guess it’s a way to double-weight swSTR, but I wanted to start very simple and only add things that I thought were useful.
Actually BB% is the component that makes the biggest difference.
Also Carrasco is 11th on my list, Hahn 158th (terrible BB%), Shark 15th, and Keuchel 23rd. Still gives some good sleeper info I think, and maybe removes a false positive in Hahn?
Er, except in this case I’m looking at the data in 2014, didn’t notice that you meant 2013 data.
BB%? So just when the at-bat ends with that pitch and a walk? Maybe zone%, but I’m not going to do at-bat enders, that’s dirty in my mind. Why would it matter if it was the last pitch or a middle pitch.
Hahn doesn’t score well in this. It’s just *this type of analysis* because his curve is sexy time.
Sure, I don’t mean to demean the value of your work, but just to add to the conversation. It’s totally cool and valid to look at guys who have great pitches even if they can’t stop walking people.
I think what’s being overlooked about BB% vs Zone% is that pitchers will sometimes intentionally throw a pitch out of the zone (in order to get the batter to chase or to set up the next pitch), but they will rarely do that when they already have a 3-ball count.
If you use Zone%, you punish pitchers for trying to get batters to chase. I know of coaches who tell young pitchers that they should NEVER throw one in the zone on 0-2 or 1-2 counts. The reason why? Because if you throw a strike, the worst likely result is it getting crushed. If you don’t the worst likely result is a 2-2 count (it could also get crushed, but that’s far less likely than if you threw a strike.)
swstr% and gb/fb is on the outcome level – velo & movement are features (or lead to outcomes) so it probably doesn’t make sense evaluating them on the same level, but obviously uber important.
Totally agree, this was sort of a first step in looking at this data. I will likely eventually separate this into two different systems that are more or less useful depending on your sample size.
Could I suggest that pop-ups be added as a component of the mythical, mystical arsenal score? Would that be relevant?
We’re totally considering this. Zimm likes PU% and Zone% as additions.
Wouldn’t including zone% unfairly penalize pitchers for those times they hit their spots out of the zone?
Separate, but related question, does zone% use pitchfx k-zones, or the ump’s call?
I pointed out that same flaw in using Zone% a couple posts up before I saw yours.
I would assume it uses the pitchfx zones, because the ump doesn’t make a call if the batter swings.
What to make of Carlos Martinez? Would this list suggest relievers have the stuff to be starters?
Should I be surprised that Chris Sale does not rank high on the list?
How about trying to add quality of contact via sOPS+ (a la http://www.fangraphs.com/fantasy/trying-to-measure-contact-management/)? Too dependent on team/park factors? Not enough sample?
I think that’s a great idea, but I’m not sure sOPS+ is the tool for the job. I wonder if the pitchf/x technology is capable of measuring things like RPMs on the ball if it was programmed right. I’d bet that has a high correlation with inducing weak contact.
More research on contact management would be a great thing for the sabermetric community.
Not sure a z-score is the right approach. That assumes a normal distribution of ‘results’ around the particular pitches in an arsenal. And I think it is safer to assume that the distribution of those results is far more of a Pareto distribution – where outliers are both more prevalent and FAR more significant than a normal distribution assumes.
Can you explain this a bit better? The underlying data is almost a perfect normal distribution.
To clarify: are you summing GB% and swstr% before you find z-scores? This assumes that a 1 percentage-point increase in GB% for a pitch is equal (however you wish to interpret this word) to a 1-pp increase in swstr%. I find that hard to believe. I would rather find z-scores for GB% and swstr% separately and then combine them to derive some sort of pitch score.
Also, 100 pitches seems too small a sample. I have no clue after how many pitches GB% and swstr% stabilize, but 100 pitches could possibly be too small to be reliable. For example, Carrasco’s slider has been filthy in even the smallest samples, but his cutter, which notched a solid swstr% this year, has been historically (and relatively) unspectacular.
Doing the z scores first and then summing.
I picked a ‘small’ pitch number on purpose. I don’t think this will be useful for pitchers with an established track record. All I want this to do is spit out some sleepers for me. So I have to use a smaller number. 100 is about 4% of a full-season starter’s arsenal, so I think that’s a decent number to call it a pitch.
I used 50 pitches in my analysis above. I think, for the purpose of calculating z-scores, it doesn’t matter how many pitches were thrown by a specific pitcher, it’s more about the pitch itself. But then when looking at the results I would be skeptical about making conclusions about pitchers with less than 150 or so pitches of a specific type thrown.
Got it. For reference, here is the confusing sentence: “I’ve summed the ground-ball and swinging strike rates for each pitch, and then found the standard deviations.”
I get the pitch count thing, but I would be interested to see it repeated at maybe a 200-pitch threshold. I should have articulated that I think you would likely see repeat offenders between the 100- and 200-pitch lists (Carrasco would certainly still be near the top). It would just, perhaps, be slightly less volatile. But really not a huge deal given the exercise performed.
nice conversation, interesting stuff. Lot of work to do, but looking forward to the output. Happy number crunching!
Interesting that guys higher up on this list tend to throw quite a few sliders. Not sure what to make of that.
Have you considered using O-swing% in this? It makes logical sense that if a guy gets more swings on pitches out of the zone the pitch it is “better”. It could give pitchers with multiple pitches an advantage as if two pitches look the same but one moves out of the zone late and get more swings
And Eno (&Jeff) – I think one comment got at this but swstr is significant more important that gb except on maybe 1 pitch …what about weighing usage and significance? I did it once by simply using the correlations of swstr and gb/fb to an expected ERA, but there’s much better ways that I’m sure you guys can run with.
Out of complete curiosity, have you run this with 2013 stat? I wonder if this type of analysis could predict injury risk. Sharper/better arsenal probably means more arm/shoulder stress.