The Great Valuation System Test: The Results
Yesterday, I shared the exciting news that my project partner Jason Bulay and I have completed The Great Valuation System Test, which involved a whopping 13 fantasy baseball player valuation systems. As usual, I feel like I could have done a slightly better job of explaining our process and goal. Essentially, we wanted to determine which valuation system most accurately converts a player’s statistical line (accounting for his position or ignoring it) into a dollar value. I eagerly awaited all the data so I could run the correlations and had my fingers crossed that the system I use, the REP method, performed well, if not the best. So it’s time to unveil the results.
The Results
The first set of results represent the correlations between the total dollar values earned and the total standings points achieved by each team.
| System | Correlation | R-squared |
|---|---|---|
| SGP – Winning Fantasy Baseball Denoms | 0.9697 | 0.9403 |
| Todd Zola REP | 0.9672 | 0.9355 |
| Jason’s Z-scores | 0.9670 | 0.9350 |
| ESPN | 0.9663 | 0.9337 |
| Razzball – Old System | 0.9633 | 0.9279 |
| Razzball – 0% Pos Adjustment | 0.9633 | 0.9279 |
| SGP – Tony Fox Denoms | 0.9630 | 0.9273 |
| Zach’s Z-scores | 0.9628 | 0.9271 |
| Last Player Picked | 0.9616 | 0.9246 |
| Razzball – 25% Pos Adjustment | 0.9615 | 0.9244 |
| Razzball – 50% Pos Adjustment | 0.9579 | 0.9175 |
| Razzball – 75% Pos Adjustment | 0.9524 | 0.9070 |
| Razzball – 100% Pos Adjustment | 0.9453 | 0.8936 |
Let’s begin interpreting these results by looking at the big picture. With correlations ranging in the mid-to-high 0.90 range, the systems do a fairly strong job of valuing players accurately. That’s a good thing. But perhaps it’s a bit of a surprise that the correlations aren’t even higher. It suggests that there may still be potential work to do to improve player valuations in order to inch our way up closer to a perfect 1.0 correlation. That of course is unlikely to ever happen, but I don’t see why we couldn’t push the correlation up into the 0.98-0.99 range.
The next overarching observation is that these valuation systems all do a remarkably similar job valuing players. The top 10 systems all sport correlations between 0.9615 and 0.9697, which is a rather small gap. We see a bit more differentiation and separation between systems when looking at the R-squared though. There is a much more clear cut grouping, suggesting that the valuation system you choose to employ does actually matter, which is not something you would come away feeling when solely looking at the correlations.
Our winner is crowned! And sure enough, it’s actually the SGP method that I abandoned over 10 years ago. The R-squared also suggests that it performed like the Mike Trout of valuation systems, with a meaningful gap between it and the next system. I was still satisfied seeing the Zola REP system finish in second, something I was nervous about considering my auction style is driven by my dollar values and I closely stick to them.
There’s something very interesting about the SGP results though. If you recall, we actually tested two different sets of denominators. The top system used the denominators from the book Winning Fantasy Baseball and those denoms weren’t even derived from leagues using the same format (the format included a second catcher). But then we find the second set of SGP results using Tony Fox’s denoms from his league sitting in a lowly seventh place. What this tells us shouldn’t be a surprise — the denominators you choose when employing the SGP method are extremely important and values could change dramatically depending on what you settle on.
I knew Tony’s denoms were from a different format (much smaller roster size) and looked far different than the other set. He convinced me to still test those values though and I’m glad I did. We now know that the SGP method does indeed work very well, but you just need to make sure your denominators are either from your own league history or from another set of leagues with the exact same format as yours.
In the caveats section of yesterday’s article , I mentioned that two users calculating values employing the same system may still get different results, and specifically mentioned that I know this was the case for the two z-score methods tested. We see from the correlation table that Jason’s z-scores ranked third, while Zach’s z-scores ranked eighth. I cannot be sure that Jason followed the exact same procedure as Zach did (he must not have), but the fact that Jason’s values performed pretty well does validate the method as being a legitimate option, though perhaps not the best. Well, at least the version of the method Jason used.
I was quite surprised by the poor performance from the Last Player Picked calculator. I used to compare my values to theirs to make sure I wasn’t totally off and their values always seemed to match up pretty well with mine, and better than other sets of values.
Aside from correlating the dollars earned to the standings points achieved, I also correlated each team’s rank within the league of the dollars earned with the place in the standings. So the correlation would be between a second place finish in the standings with a dollars earned total that ranked third. I initially didn’t feel like this would yield meaningful results as the original test above, but Jason convinced me to test it anyway. Unfortunately, the results were head-scratching. Rather than share the same correlation table as above, below is a table that compares each system’s rank when correlating dollars earned to standings points (Rank – Points column and the ranking of systems from the above table) and the dollars earned ranking with standings rank (Rank – Rank column).
| System | Rank – Points | Rank – Rank |
|---|---|---|
| SGP – Winning Fantasy Baseball Denoms | 1 | 1 |
| Todd Zola REP | 2 | 7 |
| Jason’s Z-scores | 3 | 2 |
| ESPN | 4 | 11 |
| Razzball – Old System | 5 | 4 |
| Razzball – 0% Pos Adjustment | 6 | 5 |
| SGP – Tony Fox Denoms | 7 | 6 |
| Zach’s Z-scores | 8 | 8 |
| Last Player Picked | 9 | 9 |
| Razzball – 25% Pos Adjustment | 10 | 3 |
| Razzball – 50% Pos Adjustment | 11 | 10 |
| Razzball – 75% Pos Adjustment | 12 | 12 |
| Razzball – 100% Pos Adjustment | 13 | 13 |
For the most part, the systems match exactly or are off by just one. SGP remains the king in both correlation tests. But three of the systems moved dramatically in the order in the rank correlations. Zola’s REP method dropped from second to seventh, ESPN fell from fourth to 11th and Razzball – 25% Pos Adjustment climbed from 10th to third. Jason and I couldn’t figure out what could possibly be causing these discrepancies. The only possible explanation is what concerned me to begin with that discouraged me from even calculating these correlations — that by just looking at the ranking, we’re completely ignoring the gap between teams. Because of this, it could cause funny things to happen in the valuation system correlation results. The overall correlations were also lower, which makes sense for that same reason.
The Next Steps
This test proved that the valuation system you choose does indeed matter, though admittedly your projections remain far more important. But it isn’t enough to know that on the whole, a valuation system performed better than another. Perhaps most crucial is getting individual player types right. When projection systems are evaluated, tests usually involve an RMSE (root-mean-square error) calculation. I’m no expert on that type of calculation, but it means a projection system isn’t going to perform well in an accuracy test if all its individual player forecasts are wrong, even if the average of the entire projected population turns out to be correct (I think).
Unfortunately, there’s no way that I’m aware of to test which system values each individual player most accurately. Does one system overvalue a certain position or category versus other systems, but it gets washed out in the overall results because it then undervalues another position or category? It’s very possible. I would love to test this. Any suggestions?
Ttomorrow I’ll take a look at some players with the biggest range in dollar values between the systems. We probably can’t determine which system is right, but it will be interesting to discuss and see if we could discover any trends.
Mike Podhorzer is the founder of ProjectingX IQ, an advanced fantasy baseball analytics platform that transforms projection data and in-season performance signals into actionable intelligence. He is the 2015 Fantasy Sports Writers Association Baseball Writer of the Year and three-time Tout Wars champion. He is the author of the eBook Projecting X 2.0: How to Forecast Baseball Player Performance, which teaches you how to project players yourself. Follow Mike on X@MikePodhorzer and contact him via email.
Excellent Mike. You wondered yesterday if you used a big enough sample size for this experiment. I have nothing to add to that except more questions. Is it possible you get different results in five years? Or even lessening the sample to 4 or 3 years of data? Any off the cuff thoughts on this?
Since this is the first test I’ve seen, then I’m as unsure as you. Sure it’s possible, but maybe the test was good and results stabilize quickly? I really don’t know.
We looked at some preliminary results after the first 10 leagues, and they were very consistent with the final results. So I’m pretty sure we have a good sample size at 50 leagues, and with that many teams, individual players have been combined in pretty much every way you can imagine. There could be bias from using just one season of stats, though. If we do a follow-up, we’ll use a different season.
I don’t understand your assessment of the correlations. You state that there’s little difference in the correlations but noticeable differences in the r-squared’s. Isn’t that like saying 1 is close to 2 but when you square them you can see how far apart they actually are?
Maybe. Perhaps I should have only displayed the R-squared. I don’t know, since this little experiment is totally new. Figured I’d present both metrics.
Can you find out what Jason and Zach did differently? also LPP is a a Score methods, did they actually do anything differently than LPP?
I think the differences are primarily in replacement level and positional adjustments. I looked into this a bit last year, and I found that the player pool used to calculate the averages/z-scores makes a pretty big difference. It’s not 100% clear, but I think from his explanation that Zach uses a larger player pool than I did (I’m not sure about LPP). And I’m pretty sure that he sets replacement level and does positional adjustments differently.
I actually set out to intentionally use the simplest version of z-scores that I could – sort of a Marcel system for player valuation. So I’m not sure what to make of how well it performed. All I did was calculate z-scores using the pool of “starters” (top 156 hitters, as determined using a larger sample and working through a few iterations). Then I adjusted to force the 157th hitter to zero, and translated to dollar values. I had planned to do positional adjustments so that the starters at each position would have a positive value, but it turned out that the top 156 hitters already had enough starters at each position, so the values ended up containing no positional adjustments. I weighted batting average the simplest way I could think of – multiplying each player’s AVG z-score by AB/550 (chosen purely because it seemed like a good number for a full season).
It’s an almost embarrassingly simplistic valuation system, and I’m still trying to figure out what it got right, or whether I was just lucky.
Thanks for working through such an interesting task. As I use z-scores myself, I’m curious how you are weighting AVG. I thought this weight was added to AVG when you converted AVG to xH (H – (player AB * lg avg AVG)). Are you using the extra hits method to convert AVG to a counting stat, or something else? Do you find a z-score for each player’s xH and then use your adjustment? Thanks for your time.
This is great, and I understand you are answering the more in-depth question of: specifically, which valuation system performs better.
I’m wondering whether this data can answer an even higher-level question: do valuation systems that account for position/positional scarcity (i.e., an fVARP-style system) perform better than valuation systems that are agnostic to position (i.e., systems that value all hitters equally and based only on statistics, regardless of position).
I am toying with the idea of my valuations ignoring position (and not worrying that my first 5 picks may be OF) for purposes of drafting, except for the obvious exception of ultimately filling each position.
I’ll save Mike the effort and repeat his caveat about valuation vs. draft strategy. We weren’t testing what values are the best ones to go *into* a draft with or how to treat positional scarcity during a draft. In a test like this, I would expect position-agnostic systems to do well, because all the stats count the same regardless of position, and the limited results we have were consistent with that. The draft strategy question, though, is how to approach positional scarcity to maximize the total value of your team, and there’s no way to answer that question with this type of test.
We wanted to do this. I considered submitting a Zola REP set of values without any positional adjustments, but it just doesn’t make sense. It allows for the possibility a player you must draft is valued negatively, but you must pay at least a buck for him. And that buck has to come from somewhere else. So intuitively, it makes little sense to even consider a valuation system that doesn’t account for position.
are the correlations not higher because no projection system can predict “3B X will be available 5 rounds from now, so pass on 3B Y for SS A”
That has nothing to do with it. We didn’t actually run drafts using the valuation systems. We just took rosters already drafted last year and applied the various dollar values to each team.
Does the Fangraphs Auction Calculator employ any of these systems?
Last player picked
I’ll ask what tweaks Appelman made to the LPP engine.
Mike, you are killing me! 🙂 I was told by the 2013 Tout Wars mixed draft league winner that SGP method was flawed.
This sounds too simple, but is it possible to build the rank by taking Z-scores SB, SGP BA etc. to build based on the most accurate system in a category?
I was told it was flawed too! But it’s likely all systems have some flaws, and perhaps SGP is just the least flawed. I certainly wouldn’t try merging systems and using different systems for each category.
lol I am going to have a stern word with that guy.
Wow so I literally just realized that winner was me. I could be slow at times. But yeah, I switched to REP for that very reason.
If we do a follow-up, there will almost certainly be one or more sets of values that are aggregates of other systems. As it stands, however, we only looked at total player values, so there’s no easy way to figure out which system is best in any given category.
My personal opinion is that SGP *is* flawed. It’s just that all the other systems are *more* flawed. We’re still nowhere near a system that has it all figured out.
So would you still take Billy Hamilton in the 2nd Round?
That all comes down to whether the REP method accurately valued him, and this test doesn’t tell us that. On the whole it does a good job, but that doesn’t necessarily mean each individual player and player type is correctly valued.
Rotographs articles like these should bear a disclaimer for math retards not to bother… I always just skim to the bottom to find the point.
I hate math, nor do I understand most of it, but love roto/fangraphs. How is this possible?
For your question about testing different positions against each other between the projection systems, you could use the X^2 test for goodness of fit. This would tell you whether the variations between positions are near enough to the mean to be caused by random variation. You will need to check whether or not the samples are suitable, and the large number of degrees of freedom could make it necessary to use a larger sample.
Beyond my statistical testing expertise! If you want to email me and explain, I’d be happy to look into it.
Having a bunch of similar systems compete against just 2 SGP systems means that they will be sharing all they players they ‘like’ whereas SGP will have less competition for its players. Not 100% sure this would be a huge effect but I wonder…
We didn’t draft teams using the valuation systems. We took already drafted rosters and yen applied the values from each system.
I would like to see (maybe even run myself) a different evaluation, that pits the different valuation systems against each other in a mock auction. Establish an auction nomination order, and run through the players, awarding each player to the system with the highest valuation, funds remaining and positional need. Then determine the resulting team standings using the same statistics (actual or projected) used to generate the valuations. Maybe run the simulation several times with randomized nomination order.
This method could easily incorporate pitching valuations as well. And, not to complain, because this is great work, but limiting the analysis to hitting reduces the value of the exercise. One of the perennial questions that this analysis could shed light on is the ideal hitting/pitching dollar allocation. Using this approach, one could even take a single evaluation system, and pit it against itself using different hitting/pitching splits.
Running simulations or mocks, you have to be careful to eliminate draft strategy from affecting the results. But yeah, there’s a lot more that can be tested. I don’t think the split matters as you just need to value players using a split matching your league tendency. If your league averages 69/31, it doesn’t matter if 63/37 is actually “right”. You need to use the same split as your league otherwise you’re going to go overkill on hitting or pitching at the expense of the other side.
Wouldn’t it be possible to gain an advantage by using a different split from the rest of the league, zigging when everyone else is zagging? And that’s assuming you even know the league average split; you may know what it has been for the last five years (if you keep that kind of data), but no way to know what it will be going into an auction. Pitting the same valuation method against itself with different splits might give some insight into the optimum split. If that analysis tells me 63/37 is optimal, and the league is going 69/31, I might just target 68/32 to gain the advantage without the extreme overkill effect.
Just because you use a certain split to value players doesn’t mean you’re going to leave your auction with that exact split. You could choose to spend more on on pitching than everyone else, but I would highly advise using the same split as the league.
League splits remain remarkably consistent. I have kept my own league results and they are always about the same. In Tout Wars, it’s the same thing.
i think the big differentiator that you’re missing out on, that is probably accounting for a large part of the remaining variation, is roster management – namely platooning or playing the hot hand, as well as mid-season adds, injury management, and the streaming of pitchers.
it’s the in-season management that are the difference between a good draft and a good season. and this may not simply express itself by looking at the final rosters.
That’s not what we’re testing. It is solely about the valuation system’s ability to accurately convert a player’s stat line into a dollar value, representing his worth in a fantasy league.
That’s certainly a major factor in an actual league, but this test didn’t get into that. All we did was take the rosters immediately after the draft, add up the stats and dollar values of the starters, and look at the correlations.
One of the biggest sources of variation seemed to be unbalanced teams, that won/lost one category by a wide margin. If a team had 50 SB more than the next team, the systems don’t know that 49 of those had no practical effect on the standings (except for the benefit of denying them to a competitor). Also, the correlations don’t take into account the differences between leagues. The same roster might score 50 points and rank first in one league, but only score 40 points and rank fifth in another league, just because of how players are distributed across the other teams. Obviously, that roster would have the same value in both leagues, and the differences in outcome would show up as unexplained variation.
No math to support all that, though, just my impressions from eyeballing 50 leagues worth of this.
The big advantage (maybe not that big, from the correlations) that SGP has over LPP is that it is tied to real-world tendencies.
LPP is trying to figure out the “optimal” player pool completely from scratch — without any knowledge about what fantasy players normally do. For the most part that works well, but there are a couple places where it doesn’t figure out the usual strategy.
One issue, for example, is middle relievers. LPP thinks the “optimal” pitching strategy is to push hard to maintain ERA and WHIP, whereas the normal fantasy player is willing to sacrifice a little on the ratios to get more W and K. (LPP also doesn’t know about IP minimums…)
A valuation method like SGP that is based on real-world leagues is going to do better when tested against real-world data.
Mays Copeland?? That’s interesting, but we didn’t test pitchers. Any explanation for why hitting valuations would be worse?
There’s nothing as pronounced for hitters as for pitchers, so I’m not sure where the differences lie.
I’d guess that fantasy owners value SB more highly, while LPP doesn’t think it’s worth the sacrifice in HR/RBI/BA (the profile of a lot of base-stealers). Maybe some divergence on low-BA types, but that’s also just speculation.
Either way, it’s a question of finding the balance for players who don’t contribute across the board. That’s a bigger issue for pitchers (where SP/RP contribute mostly in separate categories), but it’s not surprising if it shows up some in hitting as well.
I’m not certain who is right, but LPP and fantasy players (as a whole) occasionally disagree on where the proper balance is.
What denominators were used in the “winning” SGP strategy versus the Tony Fox denominators? If they were the same ones used as the results of the league then you have a circular problem.
I don’t know the actual denominators, but all the values were calculated before we even looked at league results, so that shouldn’t be a problem.
But then that leads to… how were they calculated? Given one SGP system worked and the other didn’t, this is highly relevant to your analysis.
They weren’t calculated, they were taken from Larry Schechter’s book and the ones Tony used for his own league.
Mike gets into this a bit in the article. The two SGP systems used denominators that were derived from observations of different real-world leagues. They had different roster formats, which explains the discrepancy. Obviously, the best way to apply SGP is going to be to use denominators that match your league settings.
Actually, the leagues we tested were somewhat different from real-world leagues, regardless of the roster settings. The test we ran was the equivalent of everybody in the league just drafting their team, setting their roster, and never logging in again. That’s not really realistic, but it was the best we could do here. I think the methodology is sound, but we can’t be 100% sure that it didn’t introduce some bias.
The problem here is you have two SGP techniques using two different denominators – neither of which that match the league settings – that give two different results. Which technique gives prescient over another? In very, very competitive leagues those numbers will generally be smaller.
Couldn’t Tony’s denominators be more “league-appropriate” for your runs, thus giving a truer result of the SGP method?
They weren’t. His was based on much smaller active rosters. The denoms that won were nearly the same rosters as the rosters we used, except 2 catchers instead of 1.
The active roster size doesn’t equate to what the final results are in a league. Separately, Larry’s denominators were AL/NL-only results I believe, which will give you different results than if it were, say, a mixed league.
The point here is that without league-appropriate denominators – the ones you should use – you have no idea how it would’ve done. Again, maybe Tony’s numbers were the correct ones.
Also, given the flaws in some of these techniques, you need to account for “real world” leagues, which use two catchers. The SGP method does not account for negative value in the worst of the player pool (i.e. backup catcher) that need to be taken, and that is one of a few flaws with the method.
Something you can’t account for is that league winners are going to tend to use these types of methods in their draft more often than the rest of the pack.
Isn’t it also hard to separate in-season management from this? Teams that draft well are also more likely to make moves to improve their team, right?
Nevermind on the second point… looked back at the first article.
The first point shouldn’t matter here either, since the values were applied to all teams after the fact. We don’t care who won the league or how they drafted their team, just whether the calculated total values for a team’s players correlate with how that team did in the standings.
Am I understanding correctly that the SGP at the top was built for a two-catcher system, and the player pools for the two Z-score entrants were of different size?
1) No idea how a two-catcher valuation would win a one-catcher league.
2) Don’t all the player pools and leagues have to match up?
The denominators from the top SGP values were from 2 catcher leagues, but values from all systems were calculated using the same 13-team rosters. Does that make sense? So all the player pools do match up.
I think both z-score entrants were calculated for the same league settings, we just made different choices on how to calculate our values, which might have involved using different player pools to calculate the standard deviations. I’m not sure, though – I don’t know the details of Zach’s methodology.
I wish we had been able to do more to ensure that all the systems were assuming the same roster/league size/positional eligibility rules. With the values we calculated ourselves it was easy, but we wanted to include the values that were available online, too. I think they were all based on the same league settings, but it’s certainly possible that we missed some discrepancies.
As Schecter himself states in his book, it’s not the denominators of SGP that matter so much; it’s the ratios between them that matter. So even if you change the denominators but kept the ratios between them, the valuation system should work the same.
I would also hazard a guess that a z-score method based on league stats (as Schecter’s are) rather than player-pool stats might perform a little better. I’ve often noticed that dollar values for SGP and z-score based on the same player pool and the same league stats are very similar.
*(as Schecter’s SGP numbers are)
I’ve coded up a number of these methods independently in preparation for NFBC this year and applied them to 5 different player projection values. As Pod mentioned the variation across player projections was much larger than the variation of the scoring method. As a practicing statistician I would look at those correlations and would say they are in all intents and purposes equivalent given the variances in league formats, owners, projection variance etc.
Thanks for running this. I think this is a worthwhile test though, it should be noted, that SGP methods (mine included) are based on actual standings point differences. The distributions seen by freezing each drafted team are not necessarily realistic.
Also, since these were 12-team MLB daily roster change drafts, teams draft with the knowledge of future roster moves – e.g., in this format, I draft tons of closers with the expectation of streaming starters so it would make my team look very poor in W/K and awesome in the other stats.
I do wonder if running the same test against the ACTUAL STANDINGS POINTS might provide any additional color. It would definitely be a better test against position adjustments. From my testing on this data, I’ve found position adjustments are neutral at best. Looks like that was the case (at least with the Razzball variations in this test).
I still think there is something
Yep. There are other holes in this analysis as well. Nice try, but I’m not convinced this wasn’t a giant waste of your time. Correlations seem way too high (not low). Pretty sure Shandler is on record saying we can get to 70% at best.
You’re not understanding the process then. 70% is for projections, not valuation systems.
Sure, you keep saying that. Garbage in garbage out here. Sorry dude the assumptions are just totally inconsistent. Your response to Enos question highlighted that.
Interesting stuff. One thing about valuation systems is that their application depends on the league settings. As far as I know, most valuation systems are based on a player’s final stat line. I play in a weekly lineup change league, and to make a perfect valuation system for a league like that you’d need to analyze production on a per week basis. For example, a platoon batter who plays 120 games in a weekly league is less valuable than a full-time player who misses 40 games due to injury even with identical final stat lines. You can sub a guy in during the 40 games missed by the full-time player, but you can’t sub anybody in for the platoon player when he is benched 1-2 games each week.
That’s why in my valuation, I add in replacement-level production pro-rated by AB for the projected missed time due to injury for some players. Tulo, for instance, gets some additional value for the replacement-level SS that will sub in for him if (when) he’s injured, whereas a playing-time risk would just get 450 ABs with no replacement added.
Can you publish a spreadsheet of the players and their dollar values in each system? I understand if it is proprietary. If not, that is a really cool data set that took a lot of time to gather. I bet some enterprising Rotographs readers could possibly expand on your research and come up with some new ideas. I certainly would like to try out some things.