The Great Valuation System Test: The Process
It all began with a comment by Jason Bulay, aka The Stranger, on a post I published in January pitting two popular snake draft strategies against each other — best player available and position scarcity:
I posted this as a comment to Cwik’s article yesterday, but I really want to see everybody put their money where their mouth is with all the draft strategy/player valuation theory. Do a draft (or multiple drafts, because small sample size) where everybody will be scored using 2014 stats (or 2015 projections if you prefer). After the draft, add up everybody’s roster using standard 5×5 scoring. See who wins, and what kind of draft strategy really gets the best roster. See who loses, and mock them without mercy.
Not that I disagree with you – your points make sense. But I think this is something we can actually test and I’d love to read about the results, so why not give it a shot?
His idea was intriguing. We’re all very familiar with the various projection systems and know that the masterminds behind them continually strive to improve them. They are also tested every year and we learn which performed best. But valuation systems get none of this treatment, as there has seemingly been little to no progress made on properly valuing players since the original systems were developed and shared.
Jason eventually followed up with an email and after many back and forth messages, we finally settled on exactly what we wanted to test and how we would accomplish our goal.
Initially, the idea was to try determining what the best draft strategy was, as that was essentially what my article was about. But with so many factors influencing what happens during a draft and what the optimal strategy is for your next pick, I figured it would be impossible to test. So instead, we eventually decided to keep things simple and develop a test of the various valuation systems. Which is the most accurate? Is it possible to determine the best system and how would we do so? Is there even a meaningful difference between the values the systems spit out or should we only concern ourselves with using the best set of projections?
As a super nerd who both forecasts player performance and then calculates dollar values, I wanted to know if the system I use is actually any good. Even if the projections were 100% accurate, if the system overvalues a category or inflates the values of middle infielders too dramatically, my auction performance would suffer. And of course, we all want to know if Billy Hamilton could truly be worth a second round pick.
The Process
From the beginning, we quickly agreed that it was only worth testing hitters. It made things much easier that way, plus I am operating under the assumption that the valuation systems are in more agreement over pitcher values than hitters.
Once we figured out what we were testing, the next step was to discuss how the testing would be performed. After a spirited discussion, I eventually convinced Jason to do things my way. We would first somehow put together or collect 2014 fantasy team rosters from as many leagues as we could. We would then use 2014 stats to generate the standings. Last, we would gather as many final 2014 dollar values as possible or calculate them ourselves, ensuring the league format these values were derived for was the same.
Once we had all our rosters and dollar values, the actual testing could take place. The testing would simply consist of correlating the dollar values earned with the standings points achieved by each team, in addition to correlation the ranking of dollars earned with the standings rank (e.g. 1st place, 2nd place, etc) for each team.
With the plan in place, we now had to gather team rosters from 2014 leagues. At first, we thought we might have to conduct mock drafts asking owners to assume 2014 stats. This would have been a nightmare. We soon came to our senses and I got in touch with Rudy Gamble of Razzball who was awesome enough to provide the team rosters for every 2014 Razzball Commenter League they ran. In total, there were 84 leagues in his file, though we only used 50 of them. Jason’s head would have exploded if I asked him to do any more than that.
The teams were composed of your standard 14 active hitters, minus a catcher. So it’s a one catcher league with 13 total hitters. I wanted to test leagues with two catchers, since the position is given the greatest position scarcity boost, which seemingly differs drastically depending on the valuation system. But we had to work with what we were given.
The Valuation Systems Tested
In total, we tested a whopping 13 sets of values. What follows is a short description of each:
Replacement Level System (REP System developed by Todd Zola of Mastersball) — This is the system I have used for over 10 years. I calculated the values used in the test myself, but given the nature of the system, two people calculating values using the same stats could still get different results.
I explained a little bit about how this system works in my Hamilton post linked to above. It adjusts player stats based on a mythical “replacement” player from each position and then calculates the player’s statistical contribution as a percentage of the category total from the positively valued player pool. It converts that total into categorical dollars earned, with that process being repeated for each category and then all summing to a player’s total worth. All categories are allocated an equal dollar amount and total player values will always equal the amount of the total league budget you choose to allocate to hitters.
Z-Scores — I used two sets of z-score values. First was Zach Sanders’ version taken directly from his end of season values post, which includes a link to his explanation of how the system works. Second was a version that my project partner Jason generated.
Last Player Picked — Values were generated at DraftBuddy, which currently hosts the LPP dollar value calculator. You can read a full explanation of how values are calculated here. In browsing through the various parts of Mays Copeland’s explanation articles, it seems that the system factors in the stats of the average player, rather than the replacement as per the REP system above, and then accounts for the standard deviation, before adjusting for position. Like REP though, it also involves a series of iterations to determine the pool of positively valued players.
ESPN — The only system without actual dollar values, we used the number from their Player Rater, which is essentially some sort of value, but on a different scale. The methodology is a black box.
Razzball Point Shares — Rudy explained the system to me as follows:
I use an SGP-based process where I subtract the average hitter or pitcher value (vs replacement hitter or pitcher value) for each player. I sum these up and then add the replacement level value to each hitter or pitcher. So if an average hitter had 0.0 Point Shares (i.e., SGPs) and the replacement level had -2.9 Point Shares, that average player is boosted to 2.9 Point Shares. The Point Shares for the modeled rostered universe are summed and then divided into the hitter and pitcher budgets respectively.
The position adjustment happens at the beginning. A position factor between 0 and 1 is set (it was at 75%, now set to 0). The initial calculation for ‘average hitter’ is PosFact * (Player Category Total – Average Category Total For All Roster-Worthy Players At That Position) + (1-PosFact)*(Player Category Total – Average Category Total For All Roster-Worthy Players). So a SS with power and no speed (e.g., Jhonny Peralta) would have higher $HR contributions because SS hit for less power BUT would take a bigger penalty on $SB since SS steal more than the average hitter.
We actually tested six different values from Razzball. We first included their posted values from their Player Rater. Then Rudy altered his system (with explanation above) and sent me the values from his new methodology with no positional adjustment and then 25%, 50%, 75% and 100% adjustments. Values from all six of those variations were included in the test.
SGP – This is the popular Standings Gain Points system explained in Art McGee’s How to Value Players for Rotisserie Baseball. This is the first system I ever used back around 2001 before stumbling upon the REP system and switching. The idea is to value players based on how many of a statistic is needed to gain a point in the standings in your league, which are termed SGP denominators. These denominators will vary by league. Adjustments for position can be made in this system as well.
We tested two sets of SGP values. One set used the denominators presented in Larry Schechter’s Winning Fantasy Baseball. However, the denominators were based on 14-man rosters with two catchers, which may differ from a set of denominators calculated in leagues with 13-man rosters. The second set of SGP values used significantly different denominators.
Both sets of SGP values were calculated by Tony Fox (@tfoxy83), author for Shandler Park, who got in touch with me via Twitter and volunteered to do the work. The second set of SGPs used denominators from his league, which uses a smaller roster size.
The Caveats
Position eligibility — Something I hadn’t thought about when we started the process ended up being a real issue. Heading into your draft/auction, it’s easy to determine a player’s position eligibility since you have your league rules and know exactly how many games are required to be played the previous year to qualify. But what about when calculating end of season values? If a player gained eligibility at a shallower position, say catcher, does he now get valued as a catcher? What if he didn’t gain eligibility until the last day of the season? Is there a cut-off date by which you should ignore new position eligibility? Even if you decide on a date, who wants to go through game logs to determine exactly when the new eligibility was gained?!
I decided to make things simple for myself and stuck with pre-season eligibility only. Unfortunately, I’m not sure how other systems handled this situation, which means that a player could be getting valued at different positions. However, Tony Fox confirmed to me that he used pre-season position eligibility as well, so at the very least, we know the SGP values use the same eligibility as the Jason z-scores and REP systems..
User quirks — In my explanation of the REP system, I mentioned that given how the system works, two people trying to run values using the same stats and position eligibility could still calculate different values. It is also possible that both the z-score and SGP methods could yield different values if the user doesn’t follow an identical process. Both Zach and Jason used the z-score method, but I could tell you right now that they did not yield identical values.
Sample Size — Our test consisted of 50 leagues. A sample of 50 plate appearances or innings pitched is tiny. Is 50 leagues too small? I don’t know, perhaps it is. We also only tested it on 2014 leagues. Maybe something quirky happened last season that benefited one valuation system over another. Probably not, but who knows.
A Big Thank You To…
Jason Bulay, who did the heavy lifting by assembling lineups from all hitters drafted for a whopping 600 teams, putting league standings together and matching the dollar values from 13 systems to the players on each team. He’ll be hiking the Pacific Crest Trail for five months beginning in mid-April and he and his wife will be blogging about their journey.
Rudy Gamble of the always informative and entertaining Razzball for providing the team rosters, of which we could not do this test without.
Tony Fox for calculating two sets of SGP values and being just as wonderful in person when I met him at the Baseball HQ First Pitch Forum in NJ.
Tomorrow, I will unveil the results.
Mike Podhorzer is the founder of ProjectingX IQ, an advanced fantasy baseball analytics platform that transforms projection data and in-season performance signals into actionable intelligence. He is the 2015 Fantasy Sports Writers Association Baseball Writer of the Year and three-time Tout Wars champion. He is the author of the eBook Projecting X 2.0: How to Forecast Baseball Player Performance, which teaches you how to project players yourself. Follow Mike on X@MikePodhorzer and contact him via email.
tomorrow!!!!!!
wow, talk about carrot and a stick….my backside hurts
I will read the follow up tomorrow……
Results tomorrow? RESULTS TOMORROW? Go to commercial.
Tomorrow! Tomorrow! It’s only a day away!
I use the last place team stats, average per player, as my replacement level with SPG. My theory is that value is anything above finishing last in your league, so a player helps if they can create value above the average of player on that last place team. I weight each category equally.
I have two other “pay for” values that match up pretty well.
One of the things we’re looking at doing as a follow-up, if we have time before I go hiking, is looking at some variations on the systems. Without going into details ahead of the big reveal tomorrow, there are still a couple of valuation questions that we weren’t able to answer, and we’re trying to think how we might answer them.
Another caveat that Mike didn’t go into is that we weren’t specifically testing what valuation system is the best one to use for your draft/auction. These are strictly end-of-season values, and the only thing we looked at was how well the total value of a team’s starters correlated with where that team finished in its league’s standings. I do think these results are applicable to preseason valuation, but there are a lot of other considerations that go into drafting a team.
SGP is flawed. You can’t assume that the worst hitter in your hitter pool is exactly $1. He needs to have negative value.
Trying to experimentally determine the truth of statements like that is exactly why we did this test.
Not sure if this test would solve that.
Did this test test for if you had too much of one category, you would take the best available hitter? Meaning, if a system took Dee Gordon, it would not be a good idea to then take Billy Hamilton?
If, by chance, one your hitters who had to be taken was a guy with only one AB last year, he would not be worth a buck. He’d be worth much less than that.
No, because that would fall under draft strategy. Like I said in the intro, that would be near impossible to test. We were only concerned with which system produced the most accurate values.
wrong. can’t have a negative score in roto, so negative player value makes no sense
WHAT A TEASE!!!!!!!! GIVE US THE RESULTS
Last Player Picked does not use average player as starting point – I’ll give you a hint which which player it does use though:
It’s called Last Player Picked!
😉
Looks cool. But there’s very little difference between the Z-Score/LPP/SGP methods. They all measure the spread within each category and compare it against the average. In fact the LPP method is I believe identical to Zach Sanders’ method.
There’s a fatal flaw that is shared by essentially every single valuation system, however, when dealing with mixed league auctions. And it is that every system fails to identify the number of players that will be bought for $1. Almost all systems project only a couple players to be sold for $1 when the real number is closer to 45. Larry Schechter has a chapter about this in his book. The way he deals with it is wonky and convoluted, and I deal with it in a more empirical (but effectively similar) way.
Valuation systems aren’t meant to mimic what happens in an auction. We want to know the inherent value of a statistical line. Which system does the best job converting the 5×5 stats into dollar values?
Converting 5×5 stats into value is important, but the purpose of a valuation system is to know how much capital to use to acquire players via the auction. Many tweaks are made to change how players are objectively valued. For instance, purposefully lowering the value of non-closer relief pitchers.
I think this is an issue that goes unnoticed. People just want the “best” valuation method without taking into account tweaks that are necessary to make it useful in a real auction, just like they want the “best” skills projection system without taking into account that they use proper playing time projections to go along with it. ZiPS would lose every single time in a raw hitting stats comparison with Steamer or the Fans, while it would do exceedingly well if proper playing time adjustments were made. As with projection comparisons, valuation systems can’t be properly compared without having all of the final tweaks that a competitive fantasy baseballer would use for a real auction.
But any tweaks is going to be league dependent and strategic in nature. Every player has some inherent value. I want to know what that is. I could then make the choice whether I want to devalue a certain category due to it flakiness like batting average and pay more for homers, which are a little more bankable. But that doesn’t mean Stanton’s stat line is inherently worth that inflated value, you’re just choosing to pay more for it.
oops, previous comment was meant to be on the main article.
With regards to LPP’s method, it does indeed compare players to the average in each category, not the last player or replacement player. The only time it uses the literal Last Player Picked is when taking the net value of the sum of z-scores. The formula to calculate the z-score for each category is: (Stat – Pool Average)/Pool Standard Deviation.
I’ve done dozens of 5×5’s using FG’s top 300 rankings, mostly on Yahoo, and by their projections, i’ve been finishing between 1st and 3rd regardless of draft position. I don’t follow the top 300 exactly as I sandbag any players I think will last another round and I tend to only take 1 SP in the first 10 rounds. Despite the popular ‘there is no position scarcity’ thing, it seems to me that fg’s top 300 really values position scarcity. It seems like every round in the first 10 the top player is a 2B.
The top 300 is very solid and takes scarcity into considertaion very well. Still no list works unless you have a strategy.
I am very intrigued by this. Just curious though, why use the 50 rosters? Doesn’t that introduce the individual biases of the drafters?
You could run a simulation so that it is purely based on 2014 stats and the given value formula. You could make a hypothetical league where each value formula represents a team. You could simulate a draft where each team (value formula) selects players based on the appropriate values, while introducing randomness in the order of selection to vary the results of each simulation. You run 50-100 simulations and generate an ensemble of results. You could then rank each value formula by the average finish in roto points. You could simulate an auction draft, but it would probably be easier to do a snake and you should get similar results.
Larry Schecter used a similar approach in his book, just looking at best player available vs. a “draft curve,” where certain positions get a bump in value.
Because we’re not looking at the actual standings in these leagues at the end of the season. We’re just using the rosters as a shortcut to get leagues to value and using final 2014 stats. Draft strategy is completely irrelevant. All we’re doing is trying to figure out which system most accurately converts a statistical line into a dollar value.
A simulation would not involve draft strategy, unless you designed it to. Players would only be selected based on their value relative to other valuation systems. It seems like the correlation method here introduces a lot of noise, when a simulation would isolate it to purely values and end results.
I’m down with a simulation if you knew of a program to run it!
We did consider drafting simulated leagues, actually. But as Mike says, you run into problems separating valuation from draft strategy. Also, we don’t have access to any software that will simulate drafts for us, and we weren’t about to do it manually.
This was actually a point that Mike and I debated quite a bit before we settled on the process we ended up using. The original articles that sparked this idea were more about draft strategy than player valuation. My original idea for this was primarily about draft strategy, and would have involved a series of mock drafts with owners using different approaches. Mike convinced me that there were just too many variables in a draft, and it would be impossible to separate the signal from the noise. And it would have taken forever. So we settled on something we could actually get done and learn something useful from.
I agree with Mike that mock drafts would be more about draft strategy than valuation. If you take the people out of it and just use logical computer code, this would reduce the variables more than using the Razzball rosters. This could be done with a scripted language (R, Python) without too much effort.
Also, we should get more stuff like this on Rotographs!!! Rotgraphs skews towards player analysis. In a 5×5 league, with highly variable stats like R,RBI,W,etc. a good valuation system is just as valuable as player analysis. (In my opinion.)
Looks cool. But there’s very little difference between the Z-Score/LPP/SGP methods. They all measure the spread within each category and compare it against the average. In fact the LPP method is I believe identical to Zach Sanders’ method.
There’s a fatal flaw that is shared by essentially every single valuation system, however, when dealing with mixed league auctions. And it is that every system fails to identify the number of players that will be bought for $1. Almost all systems project only a couple players to be sold for $1 when the real number is closer to 45. Larry Schechter has a chapter about this in his book. The way he deals with it is wonky and convoluted, and I deal with it in a more empirical (but effectively similar) way.
This is a topic I have thought quite a bit over the years. I will share a couple of quick thoughts.
If I understand your process from my quick read, I will suggest you may want to consider “drafting” new teams based on last years stats and associated dollar values (assign a valuation system to one or two teams, another valuation system to the next team or two and so on for each league you want to mock). Draft (or auction) the teams and check the results. Using last year’s teams, unless the teams had no roster moves during the season, will include a lot of noise.
You should look to the limit cases and see which systems fall apart there. This can be amusing.
Pitching is more difficult then hitting to value in roto because the player pool does not partition as well, particularly in leagues that include Holds or IPs (or other stats which alter the starter:reliever balance). It is harder to find an equilibrium for the pitcher pool than the hitter pool. Even more importantly, under most league rules, a larger percentage of value comes from pitchers not drafted into the starting line-up, but either drafted to the bench or pulled from the free agent pool. This, and not predictability or injury concern, is the main reason why more money is assigned to hitter values than pitcher values.
Actually drafting teams wouldn’t work because then you’re involving drat strategy rather than which system does the best job converting a statistical line into a dollar value. We didn’t use the actual standings from these leagues, but calculated them ourselves based on he final stats of the players drafted. That wouldn’t include any noise.
What do you mean by the limit cases?
Mike- Drafting can be “strategy free”, i.e., a team is simply assigned the highest valued player (according to that team’s valuation algorithm) that it can legally roster. Variants which take into account positional scarcity will already incorporate that aspect of strategy as part of the valuation process. This does remove the metagame (which players will be available when the next pick comes around, which stats or positions are becoming scarce, (were over-valued] etc.) but the metagame is secondary here.
If I understand you correctly, you are going to use the team rosters from last year for an arbitrary league (I assume from the start of the season, but that isn’t necessarily important) to assign stats to a team (and hence a place in the standings) and also use end of season dollar values for those players based on various valuation systems (and, I assume, calculated on the player pool drafted in that arbitrary league ) and see which of the systems correlates best with the actual standings across many leagues. Under these circumstances some of the valuation systems will be handicapped (for example, systems working off of SGP or z scores will require you to make choices for stats for which those will be calculated for which are non-obvious, systems assuming a distribution for ABs in valuing AVG (not sure if there are any, but there could be) will have an unexpected distribution if you don’t balance out players that were injured, etc . . . the league rules you are playing under may not necessarily match any set for which the models were created).
All of that said, there is value in what you are doing and I am interested in the results.
As far as (near) limit cases, I am thinking of how a system handles scenarios where everyone is projected for 1 SB or everyone except one or two or ten players are projected for one SB and the others are projected for large numbers or where someone bats 1000 in 3 ABs (or, for an SGP system, for example, the spread between the first and twelfth place teams in HRs is expected to be 10).
I still think you’d run into issues even when trying to draft strategy free. Anyway, your understanding is exactly correct. I assume when you say “actual” standings, you mean the standings generated from the drafted rosters and 2014 stats, not the standings of the actual teams in those leagues.
You’re kind of right about systems being handicapped, especially about SGP and choosing denominators. I don’t really know how z-scores work, so not sure if it’s an issue for that system too.
In terms of limit cases, I am going to look at players with the highest standard deviation in values, aka, the players the systems disagree on the value for. So that’s kind of what you’re saying, or at least one aspect of it.
By the way, you still do your own projections? I remember how well you did 2 years running in the Tango Challenge.
Mike- I was referring to standing generated from the actual stats.
I am going to read your results article next and will continue the relevant comments there.
I have not produced a full set of projections (not counting minor leveraging of the work of others) in years. What I have been doing is looking for the 50-90 players that have the widest range of potential outcomes (guys returning from bad seasons where there might be an identifiable cause [nagging injuries, usage patterns, luck], guys with playing time questions, guys with peripherals that don’t line-up, etc.) and finding value there.
Thanks for taking the time to look into this guys. Was just about to make my POS values and am freaking out a little bit, with draft day almost two weeks away!
Interest to see how you determine “accurate values”. Truth is, a player or roto statistics value changes dynamically as players are drafted out of the pool. These methods provide nice benchmarks, but for guys like schechter that supposedly calculate values down to the penny, its delusional to think there is any mathematical support to that value, especially once the players start flying off the board.
I completely disagree. Every stat line has an inherent worth, and while in-draft dynamics might suggest that at this point you should pay an inflated price for Player X, that doesn’t mean his projected stat line is suddenly worth this new amount.
“while in-draft dynamics might suggest that at this point you should pay an inflated price for Player X, that doesn’t mean his projected stat line is suddenly worth this new amount.”
how is this not a contradictory statement?
if there are 10 guys in a player pool that can steal 50 bases, assuming all else equal, the first one is as valuable as the ninth. once the 9 are gone, the 10th is a little more valuable.
Another hypothetical way to look at it – Your targeting 200 sbs. If you believe you’ve drafted 225 sbs, you likely value a pure sb player a little differently than if you had only accumulated 50.
Very simple. You could have in-draft inflation or deflation. Stars are going for under value and that saved money has to be spent somewhere, so the value of everyone else rises. For those that don’t adjust, they’re going to leave money over. Does that mean stars are now inherently worth their deflated prices just because your league decided to spend less on them? Of course not.
Maybe we’re arguing semantics here. During the draft, I agree that the 10th stolen base guy should probably be paid more (though I’m honestly not positive about that), but when valuing his stat line in a vacuum, he’s worth the same as the other 10 guys.
I think some distinctions need to be made here. Particularly:
1. between the value of a player in a league context and the value of a player to a particular team. The latter (which is, basically, the metagame [draft strategy]) will change during a draft on a restricted player pool (assuming starfish roto rules). The former only changes if players outside of the universe used to calculate values are selected. This is common, so, generally, player values change during an auction for both meanings, but not by necessity.
P.S. I think I can make a convincing argument that the 10th SB guy’s value only varies in the metagame.
“During the draft, I agree that the 10th stolen base guy should probably be paid more (though I’m honestly not positive about that), but when valuing his stat line in a vacuum, he’s worth the same as the other 10 guys.”
Then you’re agreeing with my original point. Draft isnt a vacuum. its not semantics.
You valued Hamilton based on his sbs (lets say $25, i dont even know but based on taking him in the second rd). IF there were a Billy Hamilton clone available the round after for your to take, or even better, for your next 5 picks, would you continue to take him? No. Once you acquired the first one your marginal valuation of sbs was a little lesser. So maybe you value ben revere or dee gordon lesser than someone with lesser sbs at the same point in the draft.
And thats my point. You can have a value on a guy, but it goes out the window once stats start getting removed from the pool. The method you say you use is based on a pool of players/category stats, that pool is ever changing once the draft starts.
You’re misunderstanding this test and our objective. This has nothing to do with a draft/auction. All we’re concerned about is which system most accurately converts a player’s stat line into a dollar value, telling us how much it’s worth. What other owners pay for a player or when they draft a player has no effect on that player’s or any others’ inherent value.
you misunderstood my responses. I’m talking about the efficacy of this valuation you’re attempting to derive. I’m done wasting my time with you.