THE BAT Projection System Joins FanGraphs!
What’s up, FanGraphs! My name is Derek Carty. If you’ve been into sabermetrics long enough, perhaps you recognize me from my time managing the fantasy sections for Baseball Prospectus and The Hardball Times a while back. The past half-decade or so I’ve been focused solely on daily fantasy sports, most notably with ESPN.com, Baseball Tonight, RotoGrinders, and on Twitter (@DerekCarty). I’m so pumped to be announcing today that I’m getting back involved publicly on the sabermetric side (privately, I never left!) by integrating my projection system, THE BAT, into the offerings here at FanGraphs!
If you’re into DFS, you may have heard of THE BAT. I’ve been building this system for the past eight years, and 2018 will be the fifth season it’s been available publicly, first at Fantasy Insiders and now at RotoGrinders.
So what is THE BAT, exactly? In simplest terms, it’s a projection system in the same vein as Steamer or ZiPS. Given my sabermetric pedigree, the backbone of the system has all the usual components: regression to the mean, multiple weighted seasons, aging curves, and minor league equivalencies. But it also has some very cool and unique elements.
Up until this point, THE BAT has strictly been for DFS purposes – projecting outcomes for individual games based on basic factors like hitter and pitcher and park quality but also on things that are very day-specific:
- Umpires
- Specific defensive alignments (Braves decide to play Freddie Freeman at third base all of a sudden? That’s accounted for),
- Weather factors (Temperature, sure, but also more granular things like humidity, air pressure, and special wind factors for the anomalous Wrigley Field)
- Platoon splits (Regressed, of course, using methods from The Book while also accounting for things like pitcher arm angle and pitch mix)
- And a bunch more
Player stats are backwards-adjusted to account for all of the individual circumstances they’ve faced in the past in order to gain a truer estimate of a player’s underlying talent level. On the season-level, that underlying talent is then forward-adjusted based on the circumstances the player will face for the remainder of the season. Which parks he’ll be playing in, what the weather is like in those parks at the time of year the games will take place, the quality of the opposing hitting/pitching on the schedule, etc. While I can’t give away all of the proprietary factors that feed THE BAT, a few more of the cooler ones I can divulge include:
- Pitcher role adjustments (A pitcher like Danny Salazar, who projects for innings as both a starter and a reliever, will project for different stats in each role)
- Dynamic park factors (The humidor in Chase? Accounted for. The Angels lowering their right field wall by 10 feet? Accounted for.)
- Position-based age curves (Catchers age differently than first basemen, after all)
- Exit velocity
This has led to great success for users in their DFS contests, and accuracy testing of the system has been quite good. Internal tests against other DFS systems show it to be elite, and published tests of THE BAT’s daily team and game projections consistently outperform Vegas lines, which many DFS players consider to be the pinnacle of forecasting sharpness.
Given this success, I’ve gotten a ton of requests from users who want a season-long version of THE BAT. After building out the functionality for it, I couldn’t think of a better place for this version of the product than here at FanGraphs. Not only would it be freely available as a sort of “thank you” to these users for their support over the years, but FanGraphs is also easily the go-to place for statistics, in my opinion, and it’s my hope that the addition of THE BAT proves valuable for everyone who uses it.
If you like the sound of THE BAT, check it out here at FanGraphs on the player pages, on its own sortable leaderboard pages, and in the fantasy Auction Calculator. If you find you enjoy it, I hope you’ll also consider trying out the DFS version at RotoGrinders, available for pre-sale now!
If you have any questions about THE BAT or really anything baseball-related, please feel free to drop a comment here or tweet at me (@DerekCarty). I’ll also be getting more involved on Instagram this year (@DerekTCarty), if that’s more your thing. Thanks so much to David Appelman and everyone here at FanGraphs for this opportunity. I seriously could not be more excited, and I look forward to interacting with you all this season!
Sounds like a very interesting and very complicated projection system. I have a few questions:
1. How do you decide which factors are relevant to include in your projection system? It seems like with so many factors you would be dealing with issues with “Overfitting.”
2. Is this a statistical model based approach? How are the “final” results compiled?
3. What are some main differences in how you produce your season long projections as opposed to your daily projections?
Thanks.
Thanks, good questions.
1) Overfitting isn’t really an issue because everything gets calculated organically. If a park inflates strikeouts by 5% over average and the pitcher is 10% over average but the temperature deflates it by 7% below average, well, that’s the exact impact of each and they get accounted for accordingly. Which means that things only get included if they actually matter, if the data is actually there to support it.
2) I’m not sure I quite understand this one. Could you elaborate?
3) The engine that runs both is the same. The player’s underlying talent projection is the same for both, it’s just a matter of layering on the appropriate contextual factors. For season-long, that means an aggregate of all future environments. For daily, it means whatever that’s days situation is.
Thanks for your responses.
Regarding #2, (your response to #1 gives me a clue to #2) I was trying to get at what kind program/language do you use to compile all the inputs and compute the outputs (Excel, Python, R, etc.).
But also I was wondering if how do you come up with the estimates for the factors (for example the park factors and temperature adjustment percentages you mentioned), without getting too specific do those come from a statistical model or some other data driven approach?
Ahh okay, I use SQL.
I can’t get into the nitty-gritty too much without giving away the store, but it’s very much a data-driven approach.
Looks pretty groovy. Is there historical comparison with how it compares in terms of predictive accuracy with ZIPS, Steamer, etc… ?
Unfortunately not, since this is the first year I’ve put together a full, official set of pre-season projections. It’s always just been DFS before.
first of all thank you for this amazing resource. second, if you ever do go back and put together full year projections there are probably a lot of people here (well, me at least) who would be interested to see.
Glad to hear it! And I’ll keep that in mind. The biggest issue would be figuring out the pre-season playing time and positional depth charts, since those are such a manual thing. Not sure how I’d do it for past years.
if it would help or interests you, I have the FG depth charts projections going back to 2015 that i would be happy to send to you.
I’d still need to find the time to do it, but that would make it feasible. Feel free to send those along, that’d be great. Are they from Opening Day, since they tend to change a lot?
They’d be from roughly mid-March. That’s when I tend to do my last NFBC drafts so I evaluate them based on when they’re used. Doesn’t help me much if FG gets it right after the draft window closes. Maybe not ideal for your purposes, but obviously you don’t have to use them if that doesn’t work for you. One would think Dark Lord Appleman has the authentic final versions hiding somewhere in a fangraphs database backup if he could be suitably enticed (human sacrifice, dark ritual, palette of kit kats, etc).
I’ll send you a google drive link via twitter DM. My twitter handle is the same as my FG username.
Really cool! I have some questions about it, though:
1. How do you come up with playing time estimates?
2. It seems like a very complex system. Do you have any automated testing to make sure every part is working correctly?
1. They simply use the FanGraphs Depth Chart playing time projections, with hitters’ PA adjusted a bit for team quality (i.e. the Astros will take more ABs than the White Sox because they have better hitters and will make outs at a lesser rate)
2. Partially, but not as robustly as I’d like. I’m getting there, but I do run manual checks frequently to make sure everything looks in order.
Cool, thanks for the response! Have you been using this for your own drafts this year? I’ll be watching your teams in LABR/Tout 🙂
I have! I used it entirely for Tout Wars since it was up and working with the Auction Calculator this past weekend. You can run the Tout Wars league settings and look at my team and see all the value I got 🙂
A bit less for LABR since it wasn’t and I needed to kind of MacGyver it.
Was initially very suspicious seeing Giancarlo’s 58 HR projection which seems incredibly high. Then realize that Depth Charts has the same. Stanton is projected for the best ISO since Bonds in 2004 by Steamer, ZIPS, and BAT which is sort of crazy to me (a bit suspicious cause I don’t think the park move will have nearly that large of an effect).
Yeah, the guys with the absolute monster power could be over-adjusted a bit based on park factors, but it’s pretty much all systems that see a monster year for Stanton. The swing from Miami to NY is about as big as they get for HR.
Why am I only projected for 22 steals?
I just posted a Twitter thread about this, actually.
https://twitter.com/DerekCarty/status/975849642985447426
How does your system account for Rizzo playing 2B when he isn’t really “playing” 2B? Just curious, glad to have you on board!
It basically looks at players who have played multiple positions before. When a guy plays 1B and 2B, for instance, what’s the difference in his defensive value? Obviously, he plays worse at second, and I quantify the overall effect. It’s not perfect, obviously, and some players are better suited to certain positions than others, but in this way I have a projection for every player at every position, which is better than not having one and assuming they’re league average or something.
Is there a version that does weekly projections for H2H leagues?
There isn’t, although this is a good idea, I should look into doing this.
Hey Derek,
I am a big fan of your work. Thanks, for everything that you do!
Season long question…I am attempting to generate auction values utilizing THE BAT projections, but my league uses OBP and SB-CS in place of AVE and SB for hitters and QS in place of W for pitchers. I was able to select OBP and SB-CS, but it looks like QS is not an option while using THE BAT projections. What would be the impact to pitching auction values when using QS in place of W?
Thanks!
You know, I’m not completely sure. I’ve heard this from a few people, so I’m going to add QS projections so that this isn’t an issue.
Awesome, thanks!
QS have been added
How did you account for the humidor?
Educated estimates based on the physics models (Dr. Nathan’s particularly) and what happened when Coors added the humidor
Will your system be getting regular rest-of-season updates?
Yes!
Looks like some of the categories got messed up in the Fangraphs import for pitchers. E.g., Kershaw with 703 saves and 31 strikeouts
Sorry about that. I added Quality Starts last night and it looks like the import got messed up because of it. Should hopefully be fixed soon.
Is asking for holds a bridge too far? Awesome work all in all Derek. I’m dropping these projections into my auction calculator to help me find guys that my projections may have missed. Thanks!
Holds are tougher because they’re less performance-based and more role-based, so I don’t have a great way to automate them yet. I’m told we might be able to pull them from elsewhere like we do with saves in the auction calculator for THE BAT, but we’ll have to see
And it looks like Holds have been added 🙂
hi Derek –
first let me join the chorus thanking you for the amazing system. two quick comments/questions:
1) while they are used in the auction calculator, when i go to extract the projections from the Projections tab on Fangraphs, saves comes up as zero for everyone. is it possible to correct that?
2) it seems – and this is simplistic – most projection systems are somewhat time series based with decay/again functions and regressing to the mean while others seem more reliant on multiregression analysis/system of equations to predict relevant , for example, HR/FB rates based on FB distance, pull% etc (it may be in fact that the time series uses that information and really drills down as opposed to using the more macro data) and then combing that with some equation predicting FB. Can you give me a little insight on which camp your model falls into and how – given the impressive micro data used in the daily and season long projection regarding context conditions – you use some of the micro batted ball data such as pull% or FB distance?
thanks , and again this is truly awesome.
Thanks, glad to hear you like it!
1) I’d need to ask David and see if this is possible
2) It kind of falls into both camps. It uses aging and regression to the mean for sure, but it also makes use of non-event-based data as well to sort of profile players. Spray (i.e. push/pull/etc) data, exit velocity, pitch velocity, lineup position, etc. It doesn’t have everything of this type I’d like to include (I have a list of a bunch of stuff I want to work in), but that kind of stuff is definitely in there.