Prospect Valuation Follow-Up Pt. 2

Intro
On Tuesday, I followed up my study from last week on how prospect lists translate to fantasy auction value. Last week’s research primarily identified key stratifications in career and peak value based on list placement and whether the player is primarily a pitcher or a hitter. This week, I found that the predictiveness of ranking lists relative to fantasy auction value doesn’t seem to be changing much over time.
But before crossing “prospect lists” off of my dynasty ranking model checklist, I wanted to combine these two analyses to look at how or if the ranking buckets hold when broken into more granular year groupings. This is important to my final rankings because I want whatever bucket groupings I use to be representative enough to capture a reasonable picture of the forward-looking value of a given prospect.
So, how do the buckets shake out? Well, read on.
Pitchers
Earlier in the week, I identified that the highest bucket of pitching ranks in the earliest list years was not especially correlated with overall value. Looking at the corresponding seven-year value metric I derived in Tuesday’s article, this is confirmed:
| Sample | 1 to 15 | 16 to 40 | 41+ |
|---|---|---|---|
| Full Sample (1990-2019) | $24.05 (156) | $3.79 (324) | $0.00 (845) |
| 1990-1999 | $14.15 (54) | $0.00 (100) | $0.00 (268) |
| 2000-2009 | $30.65 (51) | $8.16 (98) | $0.00 (309) |
| 2010-2019 | $24.11 (51) | $6.31 (126) | $0.00 (268) |
| 1990-1994 | $0.00 (29) | $0.00 (51) | $0.00 (136) |
| 1995-1999 | $28.15 (25) | $1.19 (49) | $0.00 (132) |
| 2000-2004 | $23.15 (26) | $3.33 (53) | $0.00 (156) |
| 2005-2009 | $46.36 (25) | $10.42 (45) | $0.00 (153) |
| 2010-2014 | $24.11 (31) | $6.93 (66) | $0.91 (133) |
| 2015-2019 | $24.89 (20) | $5.60 (60) | $0.00 (135) |
With a median value of $0 in all three buckets, the earliest prospect lists were functionally useless with regard to readily identifying pitchers which returned fantasy value. To track down the offenders:
| Player | List Year(s) | Career Value | 7-Year Post-List Value |
|---|---|---|---|
| Allen Watson | 1993 (9) | $1.46 | $1.46 |
| Arthur Rhodes | 1991 (6); 1992 (5) | $41.23 | $11.13; $11.13 |
| Ben McDonald | 1990 (2) | $75.89 | $68.60 |
| Brien Taylor | 1992 (1); 1993 (2) | $0.00 | $0.00; $0.00 |
| Chan Ho Park | 1994 (14) | $104.67 | $77.01 |
| Darren Dreifort | 1994 (11) | $25.60 | $25.60 |
| Darryl Kile | 1990 (11) | $118.58 | $24.28 |
| Frankie Rodriguez | 1992 (9) | $0.00 | $0.00 |
| James Baldwin | 1994 (8) | $14.33 | $14.33 |
| Jason Bere | 1993 (8) | $30.92 | $25.16 |
| Jose Silva | 1994 (10) | $0.00 | $0.00 |
| Kiki Jones | 1990 (6) | $0.00 | $0.00 |
| Kurt Miller | 1992 (14); 1993 (11) | $0.00 | $0.00; $0.00 |
| Mark Wohlers | 1992 (13) | $41.25 | $41.25 |
| Mike Harkey | 1990 (14) | $13.96 | $13.96 |
| Pedro Martinez | 1992 (10) | $560.08 | $211.47 |
| Roger Salkeld | 1991 (5); 1992 (3) | $0.00 | $0.00; $0.00 |
| Steve Avery | 1990 (1) | $80.54 | $80.54 |
| Steve Karsay | 1994 (12) | $21.27 | $10.49 |
| Todd Van Poppel | 1991 (1); 1992 (2); 1993 (7) | $3.58 | $0.00; $0.00; $0.00 |
| Tyrone Hill | 1993 (10) | $0.00 | $0.00 |
| Willie Banks | 1990 (13); 1991 (15) | $0.00 | $0.00; $0.00 |
As expected, other than Pedro, this is a pretty underwhelming list. Meanwhile, 10 years later, we see a relative golden age for pitching prospects within the top bucket:
| Player | List Year(s) | Career Value | 7-Year Post-List Value |
|---|---|---|---|
| Andrew Miller | 2007 (10) | $65.38 | $0.00 |
| Brett Anderson | 2009 (7) | $16.94 | $14.74 |
| Chad Billingsley | 2006 (7) | $55.52 | $55.52 |
| Clay Buchholz | 2008 (4) | $69.38 | $47.09 |
| Clayton Kershaw | 2008 (7) | $572.87 | $267.08 |
| Daisuke Matsuzaka | 2007 (1) | $30.65 | $30.65 |
| David Price | 2008 (10); 2009 (2) | $204.58 | $134.26; $171.68 |
| Félix Hernández | 2005 (2) | $265.34 | $139.99 |
| Francisco Liriano | 2006 (6) | $79.74 | $46.36 |
| Franklin Morales | 2008 (8) | $1.74 | $0.00 |
| Homer Bailey | 2007 (5); 2008 (9) | $24.76 | $22.72; $23.00 |
| Jake McGee | 2008 (15) | $37.05 | $19.10 |
| Joba Chamberlain | 2008 (3) | $10.61 | $9.01 |
| Justin Verlander | 2006 (8) | $453.34 | $206.77 |
| Madison Bumgarner | 2009 (9) | $201.91 | $142.94 |
| Matt Cain | 2005 (13); 2006 (10) | $151.37 | $112.97; $147.10 |
| Neftalí Feliz | 2009 (10) | $25.43 | $23.71 |
| Phil Hughes | 2007 (4) | $45.36 | $27.92 |
| Scott Kazmir | 2005 (7) | $74.16 | $52.36 |
| Tim Lincecum | 2007 (11) | $144.79 | $144.79 |
| Tommy Hanson | 2009 (4) | $43.87 | $43.87 |
| Trevor Cahill | 2009 (11) | $33.72 | $29.81 |
This is a superb list. There are two first-ballot Hall-of-Famers and a number of additional guys who have a reasonable case. Even in the lower end of the bucket, we see players who carved out decent careers as relievers.
Other than those two windows, though, the performance of pitchers within the top bucket has been very consistent. Based on how clearly these two buckets of time mirror each other, I’m inclined to believe they represent random talent spikes rather than any substantive statement about the lists themselves.
Calling back to the year sample split measured by career value, the break is not quite as clean:
| Sample | 1 to 15 | 16 to 40 | 41+ |
|---|---|---|---|
| Full Sample (1990-2013) | $30.65 (131) | $7.80 (249) | $0.00 (685) |
| 1990-2001 | $19.93 (66) | $2.42 (120) | $0.00 (322) |
| 2002-2013 | $54.08 (65) | $11.78 (129) | $0.00 (363) |
| 1990-1997 | $14.15 (44) | $1.05 (80) | $0.00 (214) |
| 1998-2005 | $30.47 (39) | $5.99 (83) | $0.00 (239) |
| 2006-2013 | $54.08 (48) | $11.78 (86) | $2.20 (232) |
This is intuitive though; in the dual split, the earlier bucket is dominated by the poorer performing 1990-1994 segment. Likewise for the earlier bucket in the three-way split. On the flip side, the outstanding run of pitching prospects from the late aughts is driving the high median values for the later year groupings.
With this in mind, I looked at whether the rank buckets measuring career value maintained statistical significance from each other after additionally breaking them down by year:
| Year Split | Chi-Squared | DF | P-Value |
|---|---|---|---|
| Full Sample (1990-2013) | 206.78 | 2 | 0.000 |
| 1990-2001 | 94.088 | 2 | 0.000 |
| 2002-2013 | 112.94 | 2 | 0.000 |
| 1990-1997 | 67.698 | 2 | 0.000 |
| 1998-2005 | 53.394 | 2 | 0.000 |
| 2006-2013 | 86.953 | 2 | 0.000 |
Remember, the above tests depicts whether or not there is at least one statistically significant difference between the median career value in the buckets for each year split. The tests below verify that each bucket is statistically distinct from one another.

The above two images show the results within the dual (1990 to 2001 vs. 2002 to 2013) year sample split, while the three below show the results for the triple year sample splits.

It turns out, despite the relative lack of differentiation in the earlier year buckets, they did. Given on this constellation of factors, I feel pretty good about treating pitchers based on these buckets in the dynasty model. There is still the specter of ever-changing usage patterns looming over the most recent years of these data, but I think this is mitigated somewhat by the fact that fantasy auction value each year is scaled to that year’s environment.
In other words, while it might be true that the best starting pitching prospects these days are only barely expected to qualify for the ERA title, if everyone is in that range of innings, the best players will still return the highest auction values. So, to the extent that the prospect lists are adequately identifying pitcher skill (we have no reason to believe they aren’t), we should expect that skill to likewise translate somewhat smoothly into fantasy value. This is, of course, something to closely monitor over time as more data accrue. But for now, I think that this works.
Hitters
Hitters, meanwhile, saw relatively stabler correlations over time. This tracks with what one might assume. But we all know what assuming can do, so I took a look at how the buckets behaved:
| Sample | Top 2 | 3 to 10 | 11 to 25 | 26 to 50 | 50+ |
|---|---|---|---|---|---|
| Full Sample (1990-2019) | $55.73 (45) | $22.45 (158) | $9.12 (256) | $0.25 (428) | $0.00 (788) |
| 1990-1999 | $38.62 (13) | $31.00 (54) | $13.81 (84) | $0.26 (155) | $0.00 (272) |
| 2000-2009 | $40.39 (15) | $21.77 (48) | $4.95 (97) | $0.00 (143) | $0.00 (239) |
| 2010-2019 | $101.77 (17) | $16.14 (56) | $8.61 (75) | $2.93 (130) | $0.00 (277) |
| 1990-1994 | $73.53 (4) | $22.29 (27) | $16.80 (44) | $0.00 (77) | $0.00 (132) |
| 1995-1999 | $38.62 (9) | $32.96 (27) | $12.99 (40) | $1.30 (78) | $0.00 (140) |
| 2000-2004 | $31.60 (8) | $31.37 (26) | $0.00 (44) | $0.00 (64) | $0.00 (123) |
| 2005-2009 | $64.24 (7) | $6.57 (22) | $18.85 (53) | $0.00 (79) | $0.00 (116) |
| 2010-2014 | $101.77 (7) | $15.59 (25) | $12.24 (34) | $9.21 (63) | $0.00 (141) |
| 2015-2019 | $83.36 (10) | $22.61 (31) | $7.29 (41) | $1.69 (67) | $1.82 (136) |
Well, well, well! This is significantly muddier. For one, there doesn’t seem to be meaningful distinction in the value of the top two buckets from 1995-2004. For another, the 11 to 25 bucket returned roughly three times as much value as the three to 10 bucket in the 2005-2009 sample. Likewise, there is not much separating rank three through 50 in 2010-2014, and for many of the years, there doesn’t seem to be much differentiation in the bottom two buckets. This is overall consistent with the career value measure:
| Sample | Top 2 | 3 to 10 | 11 to 25 | 26 to 50 | 50+ |
|---|---|---|---|---|---|
| Full Sample (1990-2013) | $90.04 (33) | $34.30 (121) | $17.72 (210) | $1.21 (347) | $0.00 (624) |
| 1990-2001 | $66.34 (17) | $40.42 (63) | $16.17 (101) | $2.68 (182) | $0.00 (329) |
| 2002-2013 | $105.50 (16) | $26.80 (58) | $20.47 (109) | $0.00 (165) | $0.00 (295) |
| 1990-1997 | $144.83 (9) | $47.10 (46) | $23.79 (65) | $6.35 (124) | $0.00 (218) |
| 1998-2005 | $57.19 (13) | $27.54 (41) | $4.02 (72) | $0.00 (109) | $0.00 (204) |
| 2006-2013 | $90.04 (11) | $12.98 (34) | $22.34 (73) | $0.44 (114) | $0.00 (202) |
What to do? I think this is reason enough to look into different hitter buckets. For the purposes of this exercise, I don’t think you necessarily need statistically significant breaks along every single axis; if you cut the data enough ways eventually your sample sizes will get small enough that you will inevitably manufacture problems. But given how cleanly the pitcher buckets hold over time I wanted to see if there was a way to split the hitters that stayed distinct without relying on 24 years of rankings data. I tried a few and the grouping below seemed the most promising based on median career value:
| Sample | Top 3 | 4 to 30 | 31+ |
|---|---|---|---|
| Full Sample (1990-2013) | $66.34 (53) | $22.01 (382) | $0.00 (900) |
| 1990-2001 | $61.77 (26) | $23.60 (191) | $0.00 (475) |
| 2002-2013 | $66.54 (27) | $19.50 (191) | $0.00 (425) |
| 1990-1997 | $66.34 (15) | $37.63 (131) | $0.00 (316) |
| 1998-2005 | $55.67 (20) | $3.78 (125) | $0.00 (294) |
| 2006-2013 | $57.12 (18) | $18.85 (126) | $0.00 (290) |
I liked this one for a number of reasons. First, choosing three buckets is symmetrical with the pitcher dataset. To hearken back to the credo I espoused in my debut article, I think parsimony is important. If you can achieve a similar, conceptually valid modeling outcome with fewer steps, you should. In addition, the values are fairly consistent. This again is similar to the story with pitchers. I will acknowledge that using the adjusted seven-year value metric, the picture is not quite as rosy:
| Sample | Top 3 | 4 to 30 | 31+ |
|---|---|---|---|
| Full Sample (1990-2019) | $36.34 (71) | $12.00 (471) | $0.00 (1133) |
| 1990-1999 | $22.29 (21) | $15.55 (162) | $0.00 (395) |
| 2000-2009 | $27.54 (23) | $4.50 (166) | $0.00 (353) |
| 2010-2019 | $56.67 (27) | $12.45 (143) | $0.00 (385) |
| 1990-1994 | $14.32 (8) | $17.41 (81) | $0.00 (195) |
| 1995-1999 | $40.86 (13) | $14.38 (81) | $0.00 (200) |
| 2000-2004 | $27.20 (12) | $2.49 (76) | $0.00 (177) |
| 2005-2009 | $30.78 (11) | $8.67 (90) | $0.00 (176) |
| 2010-2014 | $57.58 (12) | $13.72 (66) | $0.00 (192) |
| 2015-2019 | $56.67 (15) | $11.39 (77) | $2.42 (193) |
This is worth keeping an eye on, but isn’t something I think sinks the ship. For one, the sample sizes in the earliest bucket are naturally much smaller than those in the pitcher version of this chart. For another, the second bucket, with larger sample sizes, varies much more reasonably around the mean value. The latest bucket is clearly consistent overall.
With this in mind, I ran the statistical tests on the rank buckets for career value split out over different year groupings:
| Year Split | Chi-Squared | DF | P-Value |
|---|---|---|---|
| Full Sample (1990-2013) | 109.77 | 2 | 0.000 |
| 1990-2001 | 58.179 | 2 | 0.000 |
| 2002-2013 | 51.339 | 2 | 0.000 |
| 1990-1997 | 50.344 | 2 | 0.000 |
| 1998-2005 | 29.485 | 2 | 0.000 |
| 2006-2013 | 36.565 | 2 | 0.000 |
Consider the first bar passed. Now, to see if the individual buckets within these groupings are statistically distinct:


As with their pitcher counterparts, the above two images show the results within the dual year sample split, while the three below show the results for the triple split.


All but one pairing (Top 3 vs. 4 to 30 for 1990 to 1997) tests out. That p-value too is a whisper away from marginal significance (p=0.107). I think this is a problem with sample size – the median values shown above are very distinct, but there are only 15 hitters who ranked in the top three for those eight years. To me, these buckets pass the smell test.
Conclusion
So there you have it. Prospect rankings (at least, Baseball America’s) have generally been consistent in their ability to predict fantasy success over time; in addition, poking around the data more revealed that the typical dollar value of certain bucket rankings has likewise been fairly consistent, subject to some reasonable variation.
This exercise also proved valuable in illuminating that the pitcher buckets I unearthed last week were reliable over time, while the hitters weren’t and required a bit of reconfiguring. Pitchers, durable? Who would have thought.
Jokes aside, the end result is, in my opinion, a more robust treatment of prospect data for the dynasty model than what I had previously conducted. Stay tuned for next week when I dive into the next piece of the puzzle: aging curves.
Jonathan is a contributor for RotoGraphs. He is a Tigers fan living in Philadelphia with his wife and dog and requests that you leave your best pizza topping combinations in the comments.