The Great Valuation System Test: The Results

Yesterday, I shared the exciting news that my project partner Jason Bulay and I have completed The Great Valuation System Test, which involved a whopping 13 fantasy baseball player valuation systems. As usual, I feel like I could have done a slightly better job of explaining our process and goal. Essentially, we wanted to determine which valuation system most accurately converts a player’s statistical line (accounting for his position or ignoring it) into a dollar value. I eagerly awaited all the data so I could run the correlations and had my fingers crossed that the system I use, the REP method, performed well, if not the best. So it’s time to unveil the results.

The Results

The first set of results represent the correlations between the total dollar values earned and the total standings points achieved by each team.

System Correlation R-squared
SGP – Winning Fantasy Baseball Denoms 0.9697 0.9403
Todd Zola REP 0.9672 0.9355
Jason’s Z-scores 0.9670 0.9350
ESPN 0.9663 0.9337
Razzball – Old System 0.9633 0.9279
Razzball – 0% Pos Adjustment 0.9633 0.9279
SGP – Tony Fox Denoms 0.9630 0.9273
Zach’s Z-scores 0.9628 0.9271
Last Player Picked 0.9616 0.9246
Razzball – 25% Pos Adjustment 0.9615 0.9244
Razzball – 50% Pos Adjustment 0.9579 0.9175
Razzball – 75% Pos Adjustment 0.9524 0.9070
Razzball – 100% Pos Adjustment 0.9453 0.8936

Let’s begin interpreting these results by looking at the big picture. With correlations ranging in the mid-to-high 0.90 range, the systems do a fairly strong job of valuing players accurately. That’s a good thing. But perhaps it’s a bit of a surprise that the correlations aren’t even higher. It suggests that there may still be potential work to do to improve player valuations in order to inch our way up closer to a perfect 1.0 correlation. That of course is unlikely to ever happen, but I don’t see why we couldn’t push the correlation up into the 0.98-0.99 range.

The next overarching observation is that these valuation systems all do a remarkably similar job valuing players. The top 10 systems all sport correlations between 0.9615 and 0.9697, which is a rather small gap. We see a bit more differentiation and separation between systems when looking at the R-squared though. There is a much more clear cut grouping, suggesting that the valuation system you choose to employ does actually matter, which is not something you would come away feeling when solely looking at the correlations.

Our winner is crowned! And sure enough, it’s actually the SGP method that I abandoned over 10 years ago. The R-squared also suggests that it performed like the Mike Trout of valuation systems, with a meaningful gap between it and the next system. I was still satisfied seeing the Zola REP system finish in second, something I was nervous about considering my auction style is driven by my dollar values and I closely stick to them.

You Aren't a FanGraphs Member
It looks like you aren't yet a FanGraphs Member (or aren't logged in). We aren't mad, just disappointed.
We get it. You want to read this article. But before we let you get back to it, we'd like to point out a few of the good reasons why you should become a Member.
1. Ad Free viewing! We won't bug you with this ad, or any other.
2. Unlimited articles! Non-Members only get to read 10 free articles a month. Members never get cut off.
3. Dark mode and Classic mode!
4. Custom player page dashboards! Choose the player cards you want, in the order you want them.
5. One-click data exports! Export our projections and leaderboards for your personal projects.
6. Remove the photos on the home page! (Honestly, this doesn't sound so great to us, but some people wanted it, and we like to give our Members what they want.)
7. Even more Steamer projections! We have handedness, percentile, and context neutral projections available for Members only.
8. Get FanGraphs Walk-Off, a customized year end review! Find out exactly how you used FanGraphs this year, and how that compares to other Members. Don't be a victim of FOMO.
9. A weekly mailbag column, exclusively for Members.
10. Help support FanGraphs and our entire staff! Our Members provide us with critical resources to improve the site and deliver new features!
We hope you'll consider a Membership today, for yourself or as a gift! And we realize this has been an awfully long sales pitch, so we've also removed all the other ads in this article. We didn't want to overdo it.

There’s something very interesting about the SGP results though. If you recall, we actually tested two different sets of denominators. The top system used the denominators from the book Winning Fantasy Baseball and those denoms weren’t even derived from leagues using the same format (the format included a second catcher). But then we find the second set of SGP results using Tony Fox’s denoms from his league sitting in a lowly seventh place. What this tells us shouldn’t be a surprise — the denominators you choose when employing the SGP method are extremely important and values could change dramatically depending on what you settle on.

I knew Tony’s denoms were from a different format (much smaller roster size) and looked far different than the other set. He convinced me to still test those values though and I’m glad I did. We now know that the SGP method does indeed work very well, but you just need to make sure your denominators are either from your own league history or from another set of leagues with the exact same format as yours.

In the caveats section of yesterday’s article , I mentioned that two users calculating values employing the same system may still get different results, and specifically mentioned that I know this was the case for the two z-score methods tested. We see from the correlation table that Jason’s z-scores ranked third, while Zach’s z-scores ranked eighth. I cannot be sure that Jason followed the exact same procedure as Zach did (he must not have), but the fact that Jason’s values performed pretty well does validate the method as being a legitimate option, though perhaps not the best. Well, at least the version of the method Jason used.

I was quite surprised by the poor performance from the Last Player Picked calculator. I used to compare my values to theirs to make sure I wasn’t totally off and their values always seemed to match up pretty well with mine, and better than other sets of values.

Aside from correlating the dollars earned to the standings points achieved, I also correlated each team’s rank within the league of the dollars earned with the place in the standings. So the correlation would be between a second place finish in the standings with a dollars earned total that ranked third. I initially didn’t feel like this would yield meaningful results as the original test above, but Jason convinced me to test it anyway. Unfortunately, the results were head-scratching. Rather than share the same correlation table as above, below is a table that compares each system’s rank when correlating dollars earned to standings points (Rank – Points column and the ranking of systems from the above table) and the dollars earned ranking with standings rank (Rank – Rank column).

System Rank – Points Rank – Rank
SGP – Winning Fantasy Baseball Denoms 1 1
Todd Zola REP 2 7
Jason’s Z-scores 3 2
ESPN 4 11
Razzball – Old System 5 4
Razzball – 0% Pos Adjustment 6 5
SGP – Tony Fox Denoms 7 6
Zach’s Z-scores 8 8
Last Player Picked 9 9
Razzball – 25% Pos Adjustment 10 3
Razzball – 50% Pos Adjustment 11 10
Razzball – 75% Pos Adjustment 12 12
Razzball – 100% Pos Adjustment 13 13

For the most part, the systems match exactly or are off by just one. SGP remains the king in both correlation tests. But three of the systems moved dramatically in the order in the rank correlations. Zola’s REP method dropped from second to seventh, ESPN fell from fourth to 11th and Razzball – 25% Pos Adjustment climbed from 10th to third. Jason and I couldn’t figure out what could possibly be causing these discrepancies. The only possible explanation is what concerned me to begin with that discouraged me from even calculating these correlations — that by just looking at the ranking, we’re completely ignoring the gap between teams. Because of this, it could cause funny things to happen in the valuation system correlation results. The overall correlations were also lower, which makes sense for that same reason.

The Next Steps

This test proved that the valuation system you choose does indeed matter, though admittedly your projections remain far more important. But it isn’t enough to know that on the whole, a valuation system performed better than another. Perhaps most crucial is getting individual player types right. When projection systems are evaluated, tests usually involve an RMSE (root-mean-square error) calculation. I’m no expert on that type of calculation, but it means a projection system isn’t going to perform well in an accuracy test if all its individual player forecasts are wrong, even if the average of the entire projected population turns out to be correct (I think).

Unfortunately, there’s no way that I’m aware of to test which system values each individual player most accurately. Does one system overvalue a certain position or category versus other systems, but it gets washed out in the overall results because it then undervalues another position or category? It’s very possible. I would love to test this. Any suggestions?

Ttomorrow I’ll take a look at some players with the biggest range in dollar values between the systems. We probably can’t determine which system is right, but it will be interesting to discuss and see if we could discover any trends.





Mike Podhorzer is the founder of ProjectingX IQ, an advanced fantasy baseball analytics platform that transforms projection data and in-season performance signals into actionable intelligence. He is the 2015 Fantasy Sports Writers Association Baseball Writer of the Year and three-time Tout Wars champion. He is the author of the eBook Projecting X 2.0: How to Forecast Baseball Player Performance, which teaches you how to project players yourself. Follow Mike on X@MikePodhorzer and contact him via email.

62 Comments
Oldest
Newest Most Voted
Chicago Mark
11 years ago

Excellent Mike. You wondered yesterday if you used a big enough sample size for this experiment. I have nothing to add to that except more questions. Is it possible you get different results in five years? Or even lessening the sample to 4 or 3 years of data? Any off the cuff thoughts on this?

Jason Bulay
11 years ago
Reply to  Chicago Mark

We looked at some preliminary results after the first 10 leagues, and they were very consistent with the final results. So I’m pretty sure we have a good sample size at 50 leagues, and with that many teams, individual players have been combined in pretty much every way you can imagine. There could be bias from using just one season of stats, though. If we do a follow-up, we’ll use a different season.

Connor
11 years ago

I don’t understand your assessment of the correlations. You state that there’s little difference in the correlations but noticeable differences in the r-squared’s. Isn’t that like saying 1 is close to 2 but when you square them you can see how far apart they actually are?

draftfan
11 years ago

Can you find out what Jason and Zach did differently? also LPP is a a Score methods, did they actually do anything differently than LPP?

Jason Bulay
11 years ago
Reply to  draftfan

I think the differences are primarily in replacement level and positional adjustments. I looked into this a bit last year, and I found that the player pool used to calculate the averages/z-scores makes a pretty big difference. It’s not 100% clear, but I think from his explanation that Zach uses a larger player pool than I did (I’m not sure about LPP). And I’m pretty sure that he sets replacement level and does positional adjustments differently.

I actually set out to intentionally use the simplest version of z-scores that I could – sort of a Marcel system for player valuation. So I’m not sure what to make of how well it performed. All I did was calculate z-scores using the pool of “starters” (top 156 hitters, as determined using a larger sample and working through a few iterations). Then I adjusted to force the 157th hitter to zero, and translated to dollar values. I had planned to do positional adjustments so that the starters at each position would have a positive value, but it turned out that the top 156 hitters already had enough starters at each position, so the values ended up containing no positional adjustments. I weighted batting average the simplest way I could think of – multiplying each player’s AVG z-score by AB/550 (chosen purely because it seemed like a good number for a full season).

It’s an almost embarrassingly simplistic valuation system, and I’m still trying to figure out what it got right, or whether I was just lucky.

David Scott
11 years ago
Reply to  Jason Bulay

Thanks for working through such an interesting task. As I use z-scores myself, I’m curious how you are weighting AVG. I thought this weight was added to AVG when you converted AVG to xH (H – (player AB * lg avg AVG)). Are you using the extra hits method to convert AVG to a counting stat, or something else? Do you find a z-score for each player’s xH and then use your adjustment? Thanks for your time.

Jesse
11 years ago

This is great, and I understand you are answering the more in-depth question of: specifically, which valuation system performs better.

I’m wondering whether this data can answer an even higher-level question: do valuation systems that account for position/positional scarcity (i.e., an fVARP-style system) perform better than valuation systems that are agnostic to position (i.e., systems that value all hitters equally and based only on statistics, regardless of position).

I am toying with the idea of my valuations ignoring position (and not worrying that my first 5 picks may be OF) for purposes of drafting, except for the obvious exception of ultimately filling each position.

Jason Bulay
11 years ago
Reply to  Jesse

I’ll save Mike the effort and repeat his caveat about valuation vs. draft strategy. We weren’t testing what values are the best ones to go *into* a draft with or how to treat positional scarcity during a draft. In a test like this, I would expect position-agnostic systems to do well, because all the stats count the same regardless of position, and the limited results we have were consistent with that. The draft strategy question, though, is how to approach positional scarcity to maximize the total value of your team, and there’s no way to answer that question with this type of test.

McNulty
11 years ago

are the correlations not higher because no projection system can predict “3B X will be available 5 rounds from now, so pass on 3B Y for SS A”

Rufus T. Firefly
11 years ago

Does the Fangraphs Auction Calculator employ any of these systems?

Eno Sarris
11 years ago
Reply to  Mike Podhorzer

I’ll ask what tweaks Appelman made to the LPP engine.

Patrick
11 years ago

Mike, you are killing me! 🙂 I was told by the 2013 Tout Wars mixed draft league winner that SGP method was flawed.

This sounds too simple, but is it possible to build the rank by taking Z-scores SB, SGP BA etc. to build based on the most accurate system in a category?

Patrick
11 years ago
Reply to  Mike Podhorzer

lol I am going to have a stern word with that guy.

Jason Bulay
11 years ago
Reply to  Patrick

If we do a follow-up, there will almost certainly be one or more sets of values that are aggregates of other systems. As it stands, however, we only looked at total player values, so there’s no easy way to figure out which system is best in any given category.

My personal opinion is that SGP *is* flawed. It’s just that all the other systems are *more* flawed. We’re still nowhere near a system that has it all figured out.

YO YO MAH
11 years ago

So would you still take Billy Hamilton in the 2nd Round?

fothead
11 years ago

Rotographs articles like these should bear a disclaimer for math retards not to bother… I always just skim to the bottom to find the point.

I hate math, nor do I understand most of it, but love roto/fangraphs. How is this possible?

Anthony
11 years ago

For your question about testing different positions against each other between the projection systems, you could use the X^2 test for goodness of fit. This would tell you whether the variations between positions are near enough to the mean to be caused by random variation. You will need to check whether or not the samples are suitable, and the large number of degrees of freedom could make it necessary to use a larger sample.

Alex
11 years ago

Having a bunch of similar systems compete against just 2 SGP systems means that they will be sharing all they players they ‘like’ whereas SGP will have less competition for its players. Not 100% sure this would be a huge effect but I wonder…

Mark
11 years ago

I would like to see (maybe even run myself) a different evaluation, that pits the different valuation systems against each other in a mock auction. Establish an auction nomination order, and run through the players, awarding each player to the system with the highest valuation, funds remaining and positional need. Then determine the resulting team standings using the same statistics (actual or projected) used to generate the valuations. Maybe run the simulation several times with randomized nomination order.

This method could easily incorporate pitching valuations as well. And, not to complain, because this is great work, but limiting the analysis to hitting reduces the value of the exercise. One of the perennial questions that this analysis could shed light on is the ideal hitting/pitching dollar allocation. Using this approach, one could even take a single evaluation system, and pit it against itself using different hitting/pitching splits.

Mark
11 years ago
Reply to  Mike Podhorzer

Wouldn’t it be possible to gain an advantage by using a different split from the rest of the league, zigging when everyone else is zagging? And that’s assuming you even know the league average split; you may know what it has been for the last five years (if you keep that kind of data), but no way to know what it will be going into an auction. Pitting the same valuation method against itself with different splits might give some insight into the optimum split. If that analysis tells me 63/37 is optimal, and the league is going 69/31, I might just target 68/32 to gain the advantage without the extreme overkill effect.

metsmarathon
11 years ago

i think the big differentiator that you’re missing out on, that is probably accounting for a large part of the remaining variation, is roster management – namely platooning or playing the hot hand, as well as mid-season adds, injury management, and the streaming of pitchers.

it’s the in-season management that are the difference between a good draft and a good season. and this may not simply express itself by looking at the final rosters.

Jason Bulay
11 years ago
Reply to  metsmarathon

That’s certainly a major factor in an actual league, but this test didn’t get into that. All we did was take the rosters immediately after the draft, add up the stats and dollar values of the starters, and look at the correlations.

One of the biggest sources of variation seemed to be unbalanced teams, that won/lost one category by a wide margin. If a team had 50 SB more than the next team, the systems don’t know that 49 of those had no practical effect on the standings (except for the benefit of denying them to a competitor). Also, the correlations don’t take into account the differences between leagues. The same roster might score 50 points and rank first in one league, but only score 40 points and rank fifth in another league, just because of how players are distributed across the other teams. Obviously, that roster would have the same value in both leagues, and the differences in outcome would show up as unexplained variation.

No math to support all that, though, just my impressions from eyeballing 50 leagues worth of this.

Mays
11 years ago

The big advantage (maybe not that big, from the correlations) that SGP has over LPP is that it is tied to real-world tendencies.

LPP is trying to figure out the “optimal” player pool completely from scratch — without any knowledge about what fantasy players normally do. For the most part that works well, but there are a couple places where it doesn’t figure out the usual strategy.

One issue, for example, is middle relievers. LPP thinks the “optimal” pitching strategy is to push hard to maintain ERA and WHIP, whereas the normal fantasy player is willing to sacrifice a little on the ratios to get more W and K. (LPP also doesn’t know about IP minimums…)

A valuation method like SGP that is based on real-world leagues is going to do better when tested against real-world data.

Mays
11 years ago
Reply to  Mike Podhorzer

There’s nothing as pronounced for hitters as for pitchers, so I’m not sure where the differences lie.

I’d guess that fantasy owners value SB more highly, while LPP doesn’t think it’s worth the sacrifice in HR/RBI/BA (the profile of a lot of base-stealers). Maybe some divergence on low-BA types, but that’s also just speculation.

Either way, it’s a question of finding the balance for players who don’t contribute across the board. That’s a bigger issue for pitchers (where SP/RP contribute mostly in separate categories), but it’s not surprising if it shows up some in hitting as well.

I’m not certain who is right, but LPP and fantasy players (as a whole) occasionally disagree on where the proper balance is.

Mike D
11 years ago

What denominators were used in the “winning” SGP strategy versus the Tony Fox denominators? If they were the same ones used as the results of the league then you have a circular problem.

Jason Bulay
11 years ago
Reply to  Mike D

I don’t know the actual denominators, but all the values were calculated before we even looked at league results, so that shouldn’t be a problem.

Mike D
11 years ago
Reply to  Jason Bulay

But then that leads to… how were they calculated? Given one SGP system worked and the other didn’t, this is highly relevant to your analysis.

Jason Bulay
11 years ago
Reply to  Jason Bulay

Mike gets into this a bit in the article. The two SGP systems used denominators that were derived from observations of different real-world leagues. They had different roster formats, which explains the discrepancy. Obviously, the best way to apply SGP is going to be to use denominators that match your league settings.

Actually, the leagues we tested were somewhat different from real-world leagues, regardless of the roster settings. The test we ran was the equivalent of everybody in the league just drafting their team, setting their roster, and never logging in again. That’s not really realistic, but it was the best we could do here. I think the methodology is sound, but we can’t be 100% sure that it didn’t introduce some bias.

Mike D
11 years ago
Reply to  Jason Bulay

The problem here is you have two SGP techniques using two different denominators – neither of which that match the league settings – that give two different results. Which technique gives prescient over another? In very, very competitive leagues those numbers will generally be smaller.

Couldn’t Tony’s denominators be more “league-appropriate” for your runs, thus giving a truer result of the SGP method?

Mike D
11 years ago
Reply to  Jason Bulay

The active roster size doesn’t equate to what the final results are in a league. Separately, Larry’s denominators were AL/NL-only results I believe, which will give you different results than if it were, say, a mixed league.

The point here is that without league-appropriate denominators – the ones you should use – you have no idea how it would’ve done. Again, maybe Tony’s numbers were the correct ones.

Also, given the flaws in some of these techniques, you need to account for “real world” leagues, which use two catchers. The SGP method does not account for negative value in the worst of the player pool (i.e. backup catcher) that need to be taken, and that is one of a few flaws with the method.

Ryan BrockMember since 2025
11 years ago

Something you can’t account for is that league winners are going to tend to use these types of methods in their draft more often than the rest of the pack.

Isn’t it also hard to separate in-season management from this? Teams that draft well are also more likely to make moves to improve their team, right?

Ryan BrockMember since 2025
11 years ago
Reply to  Ryan Brock

Nevermind on the second point… looked back at the first article.

Jason Bulay
11 years ago
Reply to  Ryan Brock

The first point shouldn’t matter here either, since the values were applied to all teams after the fact. We don’t care who won the league or how they drafted their team, just whether the calculated total values for a team’s players correlate with how that team did in the standings.

Eno Sarris
11 years ago

Am I understanding correctly that the SGP at the top was built for a two-catcher system, and the player pools for the two Z-score entrants were of different size?

1) No idea how a two-catcher valuation would win a one-catcher league.
2) Don’t all the player pools and leagues have to match up?

Jason Bulay
11 years ago

I think both z-score entrants were calculated for the same league settings, we just made different choices on how to calculate our values, which might have involved using different player pools to calculate the standard deviations. I’m not sure, though – I don’t know the details of Zach’s methodology.

I wish we had been able to do more to ensure that all the systems were assuming the same roster/league size/positional eligibility rules. With the values we calculated ourselves it was easy, but we wanted to include the values that were available online, too. I think they were all based on the same league settings, but it’s certainly possible that we missed some discrepancies.

David Scott
11 years ago

As Schecter himself states in his book, it’s not the denominators of SGP that matter so much; it’s the ratios between them that matter. So even if you change the denominators but kept the ratios between them, the valuation system should work the same.

I would also hazard a guess that a z-score method based on league stats (as Schecter’s are) rather than player-pool stats might perform a little better. I’ve often noticed that dollar values for SGP and z-score based on the same player pool and the same league stats are very similar.

David Scott
11 years ago
Reply to  David Scott

*(as Schecter’s SGP numbers are)

mrrr
11 years ago

I’ve coded up a number of these methods independently in preparation for NFBC this year and applied them to 5 different player projection values. As Pod mentioned the variation across player projections was much larger than the variation of the scoring method. As a practicing statistician I would look at those correlations and would say they are in all intents and purposes equivalent given the variances in league formats, owners, projection variance etc.

Rudy Gamble
11 years ago

Thanks for running this. I think this is a worthwhile test though, it should be noted, that SGP methods (mine included) are based on actual standings point differences. The distributions seen by freezing each drafted team are not necessarily realistic.

Also, since these were 12-team MLB daily roster change drafts, teams draft with the knowledge of future roster moves – e.g., in this format, I draft tons of closers with the expectation of streaming starters so it would make my team look very poor in W/K and awesome in the other stats.

I do wonder if running the same test against the ACTUAL STANDINGS POINTS might provide any additional color. It would definitely be a better test against position adjustments. From my testing on this data, I’ve found position adjustments are neutral at best. Looks like that was the case (at least with the Razzball variations in this test).

I still think there is something

ML
11 years ago
Reply to  Rudy Gamble

Yep. There are other holes in this analysis as well. Nice try, but I’m not convinced this wasn’t a giant waste of your time. Correlations seem way too high (not low). Pretty sure Shandler is on record saying we can get to 70% at best.

ML
11 years ago
Reply to  ML

Sure, you keep saying that. Garbage in garbage out here. Sorry dude the assumptions are just totally inconsistent. Your response to Enos question highlighted that.

Luke
11 years ago

Interesting stuff. One thing about valuation systems is that their application depends on the league settings. As far as I know, most valuation systems are based on a player’s final stat line. I play in a weekly lineup change league, and to make a perfect valuation system for a league like that you’d need to analyze production on a per week basis. For example, a platoon batter who plays 120 games in a weekly league is less valuable than a full-time player who misses 40 games due to injury even with identical final stat lines. You can sub a guy in during the 40 games missed by the full-time player, but you can’t sub anybody in for the platoon player when he is benched 1-2 games each week.

TheTinDoor
11 years ago
Reply to  Luke

That’s why in my valuation, I add in replacement-level production pro-rated by AB for the projected missed time due to injury for some players. Tulo, for instance, gets some additional value for the replacement-level SS that will sub in for him if (when) he’s injured, whereas a playing-time risk would just get 450 ABs with no replacement added.

ace_hunter1982Member since 2016
11 years ago

Can you publish a spreadsheet of the players and their dollar values in each system? I understand if it is proprietary. If not, that is a really cool data set that took a lot of time to gather. I bet some enterprising Rotographs readers could possibly expand on your research and come up with some new ideas. I certainly would like to try out some things.