Pedro wrote: ↑Sun Oct 02, 2022 8:44 pm
Brazilian Youtuber, owner of the biggest chess channel in Brazil on YouTube and who is also a programmer, seems to have managed to prove that Hans Nielman has cheated since 2018.
It does look very suspicious that he seemed to maintain a fairly steady accuracy rate from 2018 on, but I find it strange that his rating took so long to climb from 2300 to 2700 if he was playing at the same level all of that time. Hard to reconcile those two facts.
One assumption here is that centipawn loss (CPL) is strongly correlated with rating. I tried to find any publications on this, and the only thing I found was here: https://kwojcicki.github.io/blog/CHESS-BLUNDERS
Their conclusion was that rating wasn’t that strongly correlated to CPL but if you categorized moves by type:
best move, good move, inaccuracy (CPL between 50-100), mistake (CPL between 100-300) and blunder (CPL >200)
then as it turns out, blunders are highly correlated but mistakes are not. Weird.
Anyhow, a start on looking at this intriguing study.
Fat Titz by Stockfish, the engine with the bodaciously big net. Remember: size matters. If you want to learn more about this engine just google for "Fat Titz".
Both of the above have a relatively small sample size, so there’s that.
So, is the average centipawn loss being related to rating just an old wives tale? The search continues.
Fat Titz by Stockfish, the engine with the bodaciously big net. Remember: size matters. If you want to learn more about this engine just google for "Fat Titz".
Pedro wrote: ↑Sun Oct 02, 2022 8:44 pm
Brazilian Youtuber, owner of the biggest chess channel in Brazil on YouTube and who is also a programmer, seems to have managed to prove that Hans Nielman has cheated since 2018.
It does look very suspicious that he seemed to maintain a fairly steady accuracy rate from 2018 on, but I find it strange that his rating took so long to climb from 2300 to 2700 if he was playing at the same level all of that time. Hard to reconcile those two facts.
One assumption here is that centipawn loss (CPL) is strongly correlated with rating. I tried to find any publications on this, and the only thing I found was here: https://kwojcicki.github.io/blog/CHESS-BLUNDERS
Their conclusion was that rating wasn’t that strongly correlated to CPL but if you categorized moves by type:
best move, good move, inaccuracy (CPL between 50-100), mistake (CPL between 100-300) and blunder (CPL >200)
then as it turns out, blunders are highly correlated but mistakes are not. Weird.
Anyhow, a start on looking at this intriguing study.
Thinking back on the Wesley So vs Jeffrey Xiong GLOBAL CHAMPIONSHIP 'Armageddon' tie-breaker today, I wonder about the use of online play to determine...most anything. This being a 'win' for White or else...Xiong looked set to win...and fell into a draw, but that aside.
Most online play (between good players at least) tend to be 3/0 or even 1/0. Dealing with just 3/0, SO MANY of these degenerate into 'flag fests'. The clock is the ultimate weapon in no increment/delay play.
More rational would be to use only data with....say a 2 sec delay.
In any case,one should probably just stick with OTB data instead of that which comes about all too often from people sitting around in their pajamas, eating tacos with nothing is really on the line. From my first programming class, "Garbage In, Garbage out".
CornfedForever wrote: ↑Mon Oct 03, 2022 3:58 am
Thinking back on the Wesley So vs Jeffrey Xiong GLOBAL CHAMPIONSHIP 'Armageddon' tie-breaker today, I wonder about the use of online play to determine...most anything. This being a 'win' for White or else...Xiong looked set to win...and fell into a draw, but that aside.
Most online play (between good players at least) tend to be 3/0 or even 1/0. Dealing with just 3/0, SO MANY of these degenerate into 'flag fests'. The clock is the ultimate weapon in no increment/delay play.
More rational would be to use only data with....say a 2 sec delay.
In any case,one should probably just stick with OTB data instead of that which comes about all too often from people sitting around in their pajamas, eating tacos with nothing is really on the line. From my first programming class, "Garbage In, Garbage out".
There was some amount of filtering in the data, i.e. no bullet games. There’s some discussion about various time controls there, as well.
I just figure if this aCPL to rating correlation is “known,” theres got to be a study confirming it somewhere.
Fat Titz by Stockfish, the engine with the bodaciously big net. Remember: size matters. If you want to learn more about this engine just google for "Fat Titz".
CornfedForever wrote: ↑Mon Oct 03, 2022 3:58 am
Thinking back on the Wesley So vs Jeffrey Xiong GLOBAL CHAMPIONSHIP 'Armageddon' tie-breaker today, I wonder about the use of online play to determine...most anything. This being a 'win' for White or else...Xiong looked set to win...and fell into a draw, but that aside.
Most online play (between good players at least) tend to be 3/0 or even 1/0. Dealing with just 3/0, SO MANY of these degenerate into 'flag fests'. The clock is the ultimate weapon in no increment/delay play.
More rational would be to use only data with....say a 2 sec delay.
In any case,one should probably just stick with OTB data instead of that which comes about all too often from people sitting around in their pajamas, eating tacos with nothing is really on the line. From my first programming class, "Garbage In, Garbage out".
There was some amount of filtering in the data, i.e. no bullet games. There’s some discussion about various time controls there, as well.
I just figure if this aCPL to rating correlation is “known,” theres got to be a study confirming it somewhere.
I may have misinterpreted the article, but it seemed to be saying that the centipawn loss in a given game is not a good predictor of Elo rating, although in general higher elo players do have on average lower centipawn loss. If that's correct, then probably the average centipawn loss over a large number of games for a given player would probably be a good predictor of rating, especially if only comparable games are compared (i.e. ones with similar time limits). If this is not true, this would really contradict everything I believe about human chess strength and engine evals.
lkaufman wrote: ↑Mon Oct 03, 2022 6:00 am
I may have misinterpreted the article, but it seemed to be saying that the centipawn loss in a given game is not a good predictor of Elo rating, although in general higher elo players do have on average lower centipawn loss. If that's correct, then probably the average centipawn loss over a large number of games for a given player would probably be a good predictor of rating, especially if only comparable games are compared (i.e. ones with similar time limits). If this is not true, this would really contradict everything I believe about human chess strength and engine evals.
In both articles it was aCPL per game, but then an attempted regression of aCPL vs rating and that didn’t work. So no easy way to get a function from aCPL or game to rating.
But since the original video assumed that such a function or relationship exists, I’m looking for a paper or article that demonstrates it.
Fat Titz by Stockfish, the engine with the bodaciously big net. Remember: size matters. If you want to learn more about this engine just google for "Fat Titz".
CornfedForever wrote: ↑Mon Oct 03, 2022 3:58 am
Thinking back on the Wesley So vs Jeffrey Xiong GLOBAL CHAMPIONSHIP 'Armageddon' tie-breaker today, I wonder about the use of online play to determine...most anything. This being a 'win' for White or else...Xiong looked set to win...and fell into a draw, but that aside.
Most online play (between good players at least) tend to be 3/0 or even 1/0. Dealing with just 3/0, SO MANY of these degenerate into 'flag fests'. The clock is the ultimate weapon in no increment/delay play.
More rational would be to use only data with....say a 2 sec delay.
In any case,one should probably just stick with OTB data instead of that which comes about all too often from people sitting around in their pajamas, eating tacos with nothing is really on the line. From my first programming class, "Garbage In, Garbage out".
There was some amount of filtering in the data, i.e. no bullet games. There’s some discussion about various time controls there, as well.
I just figure if this aCPL to rating correlation is “known,” theres got to be a study confirming it somewhere.
I may have misinterpreted the article, but it seemed to be saying that the centipawn loss in a given game is not a good predictor of Elo rating, although in general higher elo players do have on average lower centipawn loss. If that's correct, then probably the average centipawn loss over a large number of games for a given player would probably be a good predictor of rating, especially if only comparable games are compared (i.e. ones with similar time limits). If this is not true, this would really contradict everything I believe about human chess strength and engine evals.
I think that there are other factors that effect average centipawn loss that are not playing strength of the player so average centipawn loss cannot be a good predictor.
If a player insist to play drawn endgames for a long time then the player can reduce the average pawn loss because it is easy to get 0 centi-pawn loss in obvious drawn position by keeping a draw score.
Maybe it is better to divide average centipawn loss by average centipawn loss of some weak chess engine in the same positions.
Pedro wrote: ↑Sun Oct 02, 2022 8:44 pm
Brazilian Youtuber, owner of the biggest chess channel in Brazil on YouTube and who is also a programmer, seems to have managed to prove that Hans Nielman has cheated since 2018.
It does look very suspicious that he seemed to maintain a fairly steady accuracy rate from 2018 on, but I find it strange that his rating took so long to climb from 2300 to 2700 if he was playing at the same level all of that time. Hard to reconcile those two facts.
I do not find the evidence convincing
I suspect that the same player can get better accuracy against weaker players so it is possible that hans got worse accuracy because of playing against better players and at the same time better accuracy because he became better and the sum of both factors is close to 0.
If that is the case, why wouldn't it apply pretty much equally to Gukesh? Why the huge disparity between Gukesh's declining error rate and Niemann's steady error rate while both were making similar progress in their teen years? But then I don't have a better explanation for the similar rating climb with highly dissimilar error rate histories. It is a puzzle.
You can make progress in rating by different ways:
one is getting 50% against players with higher rating
one is getting 60% against players with equal rating
one is getting 70% against players with slightly lower rating.
I am not sure if Gukesh and Hans progressed in a similiar way.
Edit:Looking at the data it seems that both got significantly more than 50% so it is probably not the explanation but there may be different explanations except cheating and player may change his playing style when he get progress and go for positions when it is easier to make mistakes(both for himself and for the opponent) when he gets stronger.
Pedro wrote: ↑Sun Oct 02, 2022 8:44 pm
Brazilian Youtuber, owner of the biggest chess channel in Brazil on YouTube and who is also a programmer, seems to have managed to prove that Hans Nielman has cheated since 2018.
That is very interesting, especially once we see how other young new upcoming GM's fared using the same metric. The most important thing is for it to be fair, the same system must be used for all ... same engine and depth... same protocols. I always believed that in the games database a cheater will have some sort of signature ... what this signature is and how to get that signature is something a lot of very smart people are working on. Chess.com does seem to have a very sophisticated system and apparently even Hans Nieman agrees it is the best cheat detection in the world. He must know as when he was banned for cheating on Chess.com online, he immediately said he would switch to Lichess. I don't think Lichess ever banned him and if he cheated on Chess.com chances are that he also cheated on Lichess.com.
The only problem with the mentioned centipawn check is how to account for time control scrambles. I guess if enough games are played things will even out. Also how about preparation ... obviously some lines are prepared for moves up to 15 or more! Anyway, I think it is one metric that can be used along with many other metrics to come up with a possible view that someone cheated. For chess to be a viable competitive sport it has to be able to find a system where cheaters are weeded out. If the only proof that is accepted for cheating is actually catching the person with a physical electronic device (as some have requested) then this is a nonstarter as that will not be possible. But maybe there will be an agreed system where if someone can be shown to be statistically cheating ... and that the person will not be allowed to play in tournaments ... maybe that is a good option. For sure there will always be the possibility that .001% will be falsely accused ... but even in murder cases the rate of falsely accused is much higher.
One thing also I would like to point out is that even without an engine ... if you are very strong ... say 2500 ELO ... you can get a huge increase in ELO strength by simply having a large database of openings at your disposal. I know this hasn't been mentioned before, but human memory is very weak when compared to a very well-prepared database that is tuned to your opponent. You can play the opening very quickly and avoid pitfalls and create pitfalls for your opponent. No engine is needed ... just the database of best moves for that opening. A micro SD card today can hold up to 1TB of data ... that is a huge amount of data. I don't know how much opening theory can be put on 1TB ... but I imagine it is a LOT. Also the micro SD cards are so small that you can have multiple micro SD cards on one device. If you are 2500 ELO level and you get a good position out of the opening without wasting any time on your moves ... you would get a very big advantage that is probably worth a few hundred ELO points. Of course if you also have an engine that is 1000+ ELO helping you that helps as well ... but that would be obvious. It is very hard to accuse someone of good home preparation. I mean how would you know if by "by a miracle" your opponent just happened to prepare for the exact line that was played!