Strategy on Streaking Players: Don’t Trust the Streak
Earlier this week, on the main FanGraphs blog, we re-ran Pizza Cutter’s classic study (yes, I think it’s a legitimate classic), 525,600 minutes: how do you measure a player in a year. In it, he demonstrated just how large of a sample size you really need before you can start drawing conclusions about a batter’s skills. The answer was a lot more than I think most folks realize: you can get an idea of a hitter’s swing % and contact rate pretty quickly, but stats like OBP, SLG, and especially AVG (much less BABIP!) take 500 PA or more to provide much useful information. While I think many fantasy managers understand the need for patience, I also see a tremendous emphasis placed on small samples when I read fantasy baseball advice–especially when it comes to players on hot and cold streaks.
Is there something special about a hot or cold streak that makes it different from a typical small sample of performance? It seems like there could be, right? Even if you can’t trust a normal sample of 20 PA’s, if someone is absolutely tearing the cover off the ball–or is striking out in virtually every PA–might that not mean that he’s likely to hit particularly well (or poorly) for the next few games? After all, we see (or, at least, think we see) guys go through amazing hot streaks all the time when watching baseball, and players describe what it’s like: the game slows down, the ball looks bigger, etc.
When you actually study hot and cold streaks empirically, however, it becomes hard to find much signature to the data. I wrote about one anecdotal case last week. In that article, I referenced Tom Tango and his coauthors’ study in The Book: Playing the Percentages in Baseball, which demonstrated that there is little predictive power to a hot or cold streak. I thought that today, I might recap their study for those who haven’t read it.
The methods were simple. From 2000-2003, any five-game streak in which an individual hitter had greater than a .525 wOBA was included in a batch of “hot streakers.” And any hitter with a wOBA over 5 games of .195 or under was placed in a bucket of “cold streakers.” Note that a player could have had a 6-game streak or 7-game streak and still be included–their methods just select five consecutive hot games by a player, regardless of what happened in game 6.
Actually, what happened in game 6 is ultimately the interesting thing: if hot streaks and cold streaks are predictive, you’d expect that players in these crazy-hot and crazy-cold streaks would continue to be at least somewhat hot or somewhat cold in their next games. Here’s what happened:
Hot streakers
During the streak: .587 wOBA
Expected wOBA (just a 3-year average for the players): .365
Actual wOBA in game after streak: .369
Actual wOBA in five games after streak: .369
Cold streakers
During the streak: .151 wOBA
Expected wOBA (3-year average again): .336
Actual wOBA in first game after streak: .330
Actual wOBA in five games after streak: .332
So, is there anything predictive about hot and cold streaks? Yes, there is…but it’s slight. If you pick up an insanely hot player, or bench an insanely cold player, you get about 5 points of wOBA (a bit less than 5 points of OBP) of predictive power. 5 measley points! This is quite frankly worthless in terms of making fantasy baseball decisions. You’ll do far better relying on a player’s overall talent than trying to catch a hot streak or avoid a cold streak–because chances are, by the time you react, it’s already over.
There’s one caveat I’ll add, however, from a fantasy perspective: predictive or not, real baseball managers look at hot and cold streaks. And even though they are generally not predictive, managers will give a guy more PT when he’s “hot.” The perfect example of this is Jed Lowrie, who had an insanely good week that, apparently, at least for now, has resulted in him securing the Boston starting SS job. Does Jed’s hot streak mean anything about his performance over this coming week that is different than if he’d hit like a normal human? Probably not–except that he now has a job (and maybe we bump up our estimate of his true talent level a couple of zots). Of course, a bad cold streak could reverse all of that pretty quickly…
For more on this topic, I’d highly recommend The Book. A major portion of what’s in there is directly applicable to fantasy, so it’s worth the read for that reason alone.
Justin is a lifelong Reds fan, and first played fantasy baseball on Prodigy with a 2400 baud modem. His favorite Excel function is the vlookup(). You can find him on twitter @jinazreds, even though he no longer lives in AZ.
On cold guys, especially pitchers, you have to worry about a possible undisclosed injury.
The story is actually different for pitchers. Hot starters actually can be expected to stay hot (though to a lesser degree) in the start after their hot steak. Surprisingly, cold streaks showed a weaker effect–I would have expected it to run the opposite direction given the injury issue you mentioned. -j
Well, just as AVG stabilizes much later than Swing%, couldn’t we look at a player who’s getting on base at a ridiculous rate, or hitting the ball unusually hard, rather than just looking at wOBA? The wOBA study is a good argument against managers starting and sitting players based on streaks, but unless you’re in a wOBA league, I’m less interested in how productive a player is in real life than how productive he is in whatever set of stats I use. I’d also like to see the size of the “streak” expanded; it’s pretty obvious that a week of work shouldn’t change how you view a player much, but what if they’ve been doing something for 10 games? 20?
Clarifying: Pizza Cutter’s sample size work is beyond fantastic for looking at true talent level, and nothing I said above was intended to contradict him on players’ actual ability as it pertains to sample size. I’m only talking about “streaks.”
Agreed that this only applies to streaks in iverall production, i.e. WOBA or overall slash line success as an extension. But that’s usually what people are looking at when talking about streaks.
They did look at 7-game streaks and saw the same sized effects. If you want longer, my question is at what point does something cease to be a streak and become exactly the kind of thing that Pizza’s study was looking at. If you’re looking at the first 18 games this season, I think people would tend to think it’s not just someone playing “hot” but rather what might be an actual change in talent level. And at that point, Pizza’s study would be appropriate…
-j
Justin, thanks for the response. I think streaks and what Pizza Cutter did are completely different. Pizza Cutter determined how many PAs it took before deviations in certain stats could be said to have a 50% chance of reflecting a change in underlying talent as opposed to just normal variation. The question with streaks is, if I do something extra well/badly for 5/7/10 games, will I perform better/worse in my next 1/3/5 games than my true talent level.
Pizza Cutter discovered that we need a lot of PA to expect that an OBP different from past performance likely indicates a change in a player’s true talent at getting on base. But no one’s investigated whether or not 50 PA of a higher/lower OBP has any predictive value on the next 20 PA.
Again, though, how likely is it that whatever “zen” a hitter experiences during a hot streak will last a full 10 or 20 games? Given that there’s no improvement between 5 and 7 games, I’m skeptical that 10 games will make a difference. 20 games…again, I’m not sure that can really be considered a hot streak at that point, because it’s the better part of a month.
There is legitimate question about whether hot and cold streaks really exist, meaning that there is something different about hitters going through a hot or cold streak that is causing the streak. If you just allow a simulation to run and have equal probabilities of getting a hit throughout, you will see “hot streaks” and “cold streaks” in the data. The streaks aren’t real, though–they’re just a false “pattern” that we think we’re picking up on. I think some hot and especially cold streaks are most certainly real, especially when you’re talking about players playing through injuries (or, perhaps, being unusually healthy for a time). But I think the majority are likely just random “good” or “bad” runs, much like you’d see if you flipped a coin or rolled the dice 100 times. And I’m not at all confident that we can tell the difference between the two.
That might make for an interesting little community event: choose the “real” streaks vs the “fake” streaks. Bonifide or Bonifacio, but crowdsourced, for you Fantasy Focus listeners out there.
-j
But Justin, 20 games is what, 80-100 PA? How many of Pizza Cutter’s stats showed up as statistically significant in that few PAs? Not many. OBP took many times that before you could start adjusting expectations going forward. So why are you saying that if someone gets on base a ton for a month we shouldn’t consider it a streak? Pizza Cutter demonstrated that it’s pretty worthless for predicting how a player would do in other 100 PA sets, but, again, no one has investigated how effective it is in predicting the NEXT N plate appearances.
I’m not arguing that streaks exist, I’m arguing that there needs to be more tests than whether wOBA over 5 games can predict the 6th or next 5 games. If a player mashes for 5 games, I’m not going to assume his next 5 games he’ll continue to play at a hall-of-famer level, and his manager shouldn’t “keep the hot hand” in. But if a player walks at twice his typical rate for 45PA, I wonder if he won’t continue to walk more frequently than his typical level in the next 20PA. Pizza Cutter demonstrated that those 30PA are not useful to predict his walk rate over a long period of time, but I want to know if it’s useful for his next set of games.
If this post were on the main blog, my argument would be pointless since wOBA is production and that’s what matters. But we’re at Rotographs, and actual production is irrelevant compared to specific stats.
If the question is whether this applies to other stats–non-overall production stats–I agree that this remains an open question.
If it’s a points league, though, which is what I tend to play, the overall production is what matters. If you’re interested in category leagues, and are looking to catch a good HR streak (which is the only application I can really see for this), then maybe it is something that could be of interest. I’d be skeptical given this study, but I agree there aren’t data to answer the question.
-j
@byron,
The problem with you extending the definitions of a “hot” streak to, say, 15 games is that you also end up trying to play the “hot” player much longer. The rationale becomes, “hey that streak was over 15 games, maybe I should give him more time to prove that the streak is over.” This becomes even more problematic if you truly are trying to catch an unlikely break, i.e. a player whose career numbers suggest a much lower talent level. You end up starting a chump for 10-15 games.
Phillie, you’re probably right, but we can’t be sure until someone does the research I’ve been talking about. It is very possible that there is no such thing as a streak useful for fantasy purposes; my skepticism is that that’s already been proven.
So if a player had a 10 game “hot” or “cold” streak, would the analysis consider it “two” streaks and evaluate both games 6-10 and 11-15 for the 5-game-following streak-end analysis? IMO the game 6 stat is meaningless from start – one game so statistically insignificant. Also I’m assuming the analysis looked for 5-game averages, not for example 5 consecutive games in which wOBA was over/under x for all 5 games?
Not positive on the first question, but I think, yes, that’s two 5-game streaks. They did allow more than one streak per player, which you can argue about whether it’s appropriate, but if I remember right they checked and it didn’t affect the outcome.
Re: 1 game–the sample size is huge here, well over 1000 streaks, so it’s not “just” one game–it’s one game per streak. But they also reported data for the five games that came after the streak and saw the same lack of an effect.
Finally, yes, it’s combined stats for all five games.
-j
I should add–the sample includes 600+ players making those 1000+ streaks.
Or is it six 5-game streaks (games 1-5; 2-6; 3-7; 4-8; 5-9; 6-10)? Or in the simpler case, if there’s a 6-game hot streak there’s no reason to look at the effects of games 1 through 5, but not those of games 2 through 6.
My first reaction was to say that they wouldn’t have double-counted games, but it turns out that they did. I’m not sure, but I think that bothers me that they did. It probably doesn’t affect the results, though!
-j
I think most smart owners understand that hot streaks don’t keep up. But what they’re doing is trying to find the diamond in the rough…the Jose Bautista, whose streak turned into an underlying misjudgment of his original value.
Beneath that extra .004 of wOBA there is a batch of players who highly exceeded their expected wOBA. Your study just blended all the samples, which in turn diluted such cases.
So I think in this regard, it would be useful to show the # of hitters who performed way above expectation (say, .100 higher), and then show that # as a % of those observed.
What do you think?
Agree that it would be interesting to know. I think the key question, however, is whether the damage you do by playing crappy players with recent success is overcome by the rare Jose Bautista that you catch. Jose Bautista’s don’t come around every year…
All of this said, I do sometimes use streaks myself. In my 20-team league, with Mauer on the DL, I had to choose this week between Jonathan Lucroy and Josh Thole. Their projections are basically identical (very light edge to Lucroy, not enough to mean anything), and both are getting plenty of PT. But Lucroy’s hot, and Thole’s not. So, with nothing else to really go by, I went with Lucroy. I’m not expecting much from him, but hopefully he won’t go into a tailspin next week…
-j
Fire for effect!!
Maybe this is wishful thinking, but I’d love to know if the same conclusions apply to players with very good projected wOBA’s as players with very bad projected wOBA’s. Maybe there’s something to the “players play better the first time around the league” theory?