2019 xADP, New and Improved
About a month ago, I published a post that predicted 2019 ADP (“xADP”) values using eight years’ worth of average draft position (ADP) data from the National Fantasy Baseball Championship (NFBC) and end-of-season (EOS) values from Razzball. The model was pretty good — it explained nearly 60 percent of the data’s variance (adjusted r2 = 0.59), which is pretty dang good. It felt unfulfilled, though; it accounted for some players but not others — namely, breakout rookies who were completely off the radar the previous season and top prospects who had yet to debut.
I took some time (really, a lot of time) to clean up my data to see how much it would improve my model, if at all:
- Originally, my data set did not account for players who were not drafted (aka had no ADP value) but made an impact in 2018 (think Juan Soto). Conversely, my data did account for players who were drafted but made no impact in 2018 (think, uh, Troy Tulowitzki, I guess). It was kind of like addressing a Type I error but ignoring a Type II error (or the other way around? I don’t know). I took painstaking care to fill in these holes.
- I took equally painstaking care to ensure all player names were consistent — no “Nick Castellanos”/”Nicholas Castellanos” mismatches that might pollute the analysis. Odds are, there are a couple of players I missed, but having spent hours poring over the data, I feel confident that the issue is no longer pervasive.
- I added ages! They make a small impact, most meaningful to players at the extremes, such as the very young (think Ronald Acuna) and the very old (think Nelson Cruz).
- Lastly, a theoretical and methodological adjustment: I forced negative ADP values to $0. I wanted the model to reflect an actual draft, in which players are never bought at auction for negative dollars — rather, their values converge on zero. It’s important to note here that a player can still end the season with negative value based on the concept of replacement level. Accordingly, only negative ADP values, and not negative EOS values, were forced zero.
Fortunately, the extra work was worth it: the model boasts an adjusted r2 of 0.75 (with ages; 0.73 without). That’s a massive improvement, and it can be attributed almost entirely to the slight (but profound) change in the model specification.
The model still has its shortcomings. It doesn’t know who the top prospects are, who won and lost the closer role last year, or which seasons were shortened by injury or delayed call-ups. The most notable discrepancies will be among players who fit the latter-most category (shortened season by delayed call-up); Acuna and Ozzie Albies rank 35th and 41st, respectively, which is probably a bit low but not obscenely so. On the other hand, Walker Buehler and Adalberto Mondesi are ranked 165th and 187th, and I’ll bet the house that they each go at least 100 picks earlier.
Without further ado, here’s an updated top 500. I’ll probably rely on it gently during these early NFBC Draft Champions and Fantrax best ball drafts as well as while navigating offseason ottoneu trades. The table is sortable — just click the column headers. Go crazy!
Gabriel Moya over Jesse Biddle? That invalidates everything.
noooooOOOO
Can’t imagine there will be too many leagues where Juan Soto goes 147.
Sure, it’s conceivable he ends up producing only that sort of value, but he’s going to get drafted above that everywhere.
As mentioned in the post, the model will miss hardest on guys like Soto. Just very little precedent for their trajectories.
Acuna, with a few more MLB PA’s and slightly better stats than Soto is ranked over 100 picks higher. Something is wrong with your correlation estimates … real life results have to be a significant factor in them.
I’ll forgive you for probably not ready the original xADP post, or maybe even this post, and also maybe even my reply to the comment to which you also replied.
Soto has very few comps in terms of coming out of nowhere to make a huge impact. The model doesn’t understand he’s a top prospect; all it sees is he did not play in 2017, went completely undrafted in 2018, and then earned a huge profit. Most recent performance gets the heaviest weight, but the model also weights ADP ($0) and prior-year earnings (negative, because he didn’t play), and it makes it look like he was a fluke.
Acuna, on the other hand, was a top-100 pick in 2018. The model sees that there was an expectation for him to succeed.
No further ADO, not adieu
I said ado!
Is there a way to convert these to rough dollar values?
Ready for something mildly complicated?
$$$ = [-ln(rank) + 6] * 9.7
So, for #1 overall: [-ln(1) + 6] * 9.7 = ~$58 for a 15-team league with 27 roster spots. That should work. Maybe I goofed. Sorry I’m at work. If that’s off, holler back.
If you’re an R user, I use the following estimate (np is number of players, nt is number of teams, cap is cap per team), which should give a very similar curve: np=310; nt = 10; cap = 350; (log(np)-log(seq(np))) / sum(log(np)-log(seq(np))) * nt * cap + 1
https://www.fangraphs.com/fantasy/adp-to-auction-values-process/
Kershaw before every pitcher except Scherzer is somewhat befuddling. Intuitively it makes sense I suppose, and I love him, but I would never take him above Sale, DeGrom, Verlander or Kluber at this point.
Kershaw as the #3 pick in 2017 and the #4 pick in 2018 is buoying his value. The drop-off in 2018 (in terms of EOS value) was severe but not enough to chip away at what is perceived to be high draft value in the past. Scherzer’s highest ADP ever was last year — at 11th. That gap between him and Kershaw is what has kept Kershaw so close (and ahead of everyone else).
…or Bauer or Cole or Snell… Kershaw’s injury history is impossible to ignore at this point: 3 years between 150-175 innings
Great stuff! Between this and #2earlymock experts adp I think we have very good expectations for what ADP will look like this year. It’d be interesting to compare the two, qualitatively or quantitatively. Also, do you think using value per plate appearance would be more accurate then EOS value on razzball? I’d expect it would as long as there was a high enough minimum PAs (200-400). Guys like Mondesi, Buehler, and Soto would rank much higher.
#2earlymock: (https://docs.google.com/spreadsheets/d/1Y_4czFhvl97m1FgRGd3Fg0o_13pp2ZQYydYEo0MoD6I/edit#gid=604608158)
You have J-Up a lot higher than Casty personally?
It would be helpful to an extent but would act really wildly for guys with very small PA or IP. But I do think something like that would help; specifically, these three variables: value, playing time, and the interaction between the two. That would resolve extremes at both ends of each and effectively make the relationship nonlinear. It’s something I’ve been thinking about but you helped bring some clarity to it!
Awesome 🙂 I agree the interaction variable is a good idea!
Surely Danny Jansen should be in the top 500 right?
Yeah, but having barely played in 2018, the mode has no idea he has value. Just further proof that a top prospect variable would be helpful, and also, now that I’m thinking about it, a 2019 playing time projection (but the question would become, according to which projextion system?).
I understand the model does not see everything perfectly (young guys etc) but the bottom line is until it does, it will miss badly on a lot of players. Acuna will be a 1st round pick in nearly every draft this year, and Soto a 2nd round pick. Altuve will never go # 5. Most drafts he will be between 10-20. And of course Walker Buehler will be between 15 and 30. (Not 165) Blake Snell will never last until mid 50s either. This is a fun model to look at and thanks for doing it, but it is way too low on the young stars. As a group of drafters, we clearly go hard for young guys with good upside.
You’re probably right. The sad part about the model is it’s going to be most accurate on the boring guys who no one has a bullish or contentious valuation for. The guys who will be hot topics in any given year will probably be the ones where there’s some vocal disagreement, aka, inherent variance. And, yes, it appears to still not appreciate our shiny new toy bias.
Interesting reading. I read the caveats in both articles, so no questions about Vlad jr’s spot. Curious why it’s quite so low on Lourdes Gurriel, is that a multiple names issue? To pick a name at semi-random, Gurriel’s over 100 spots behind Guzman, even though I think both were practically undrafted in 2018, then got regular playing time. Gurriel performed better than Guzman did, although only in about half the games. Razzball’s rater has Guzman at 406 and Gurriel 586, so I assume that’s also the cause of the difference in xadp. I wonder if incorporating Razzball’s $/game rank rather than just overall rank might improve it for those who didn’t play full seasons? That one has Gurriel at 179 & Guzman at 329, which is closer to what I’d expect their relative adps to be.
A couple of other potential weird ones. What’s causing Zach Britton to miss the top 500 altogether? I assume he has some highish ADPs from a few years ago, he’d have been drafted a bit last year to stash on the DL as a potential 2nd half closer. Other speculative 2018 guys like Brach & Bedrosian who finished just behind Britton on the player-rater did make xadp’s top 500.
And finally, why does it see Margot & Gleyber Torres so similarly? Both are around 150 with age included, 175 without. That seems off for both of them. If I had to guess, Torres is probably getting a bit less credit than he should because he only played 80% of the season. Then it’s seeing Margot as ~150th overall pick last year, his low position on the player rater this year, but not docking him much for it. Most other guys who do so poorly after being taken that high have missed significant time and therefore keep most of their adp the following year. e.g. Josh Donaldson being a top 30 pick last year and xadp keeping him inside the top 100, Cano being a top 70 pick last year and now ~170. Or Adam Eaton, who like Margot was drafted around ~150 last year, like Margot ranked at ~360 on the player rater, and so like Margot is pegged by xadp to go around 150-175 this year. But because Margot finished 360th due to a full season of being poor, and Eaton did it due to half a season injured and half a season being decent, their projections are quite different and I’d expect their adp will be too. Which is reflected in their $/g positions on razzball’s rater, ~125 vs 363.
Who are the guys at #363 & #400 on the without age list, that drop right out of the top 500 once their age is considered?
Where is the fat shirtless Colon guy when you need him for a timely post?
Really useful approach. Being in a single league roto league, would the ranking still hold if other league players are removed? And if so, could you republish the table with team abbreviations added?
Is Shohei Ohtani listed twice because you have him ranked separately as hitter and pitcher?
Oh, weird. No, I think that’s just a weird thing that happened when I made the table. I actually think his name replaced someone else who apparently isn’t that important. His actual xADP is the first one (mid-50s or whatever).