Since A.I came into existence, my little project got debugged much easier. Is that a shame ? I'll admit something even deeper : When i asked it to analyse my code, it just laughed in my faceGabor Szots wrote: ↑Sun Jul 12, 2026 12:21 pm Accepting that AI engines are the future, I have questions from a tester's point of view.
To begin with, I have been living in the assumption that the goal of the rating lists is to compare the skills of the authors. Now with AI generated engines that goal seems obsolete because there is no author in the traditional sense.
So:
1. If I decide to test an AI engine whom should I credit? Claude? Some other AI program?
2. Which AI engine should I test? They are all gathering the same knowledge, as I see it there should be no difference between them.
3. Has copyright and licencing become obsolete?
Testing AI engines
Moderator: Ras
-
urbanmusic
- Posts: 16
- Joined: Tue Dec 12, 2023 11:36 pm
- Full name: Bruno Santos
Re: Testing AI engines
-
dkappe
- Posts: 1636
- Joined: Tue Aug 21, 2018 7:52 pm
- Full name: Dietrich Kappe
Re: Testing AI engines
There are certainly vibe coded engines, but using an AI is not necessarily the same as vibe coding.
For example, if you practice TDD, you can write tests and instrument search in a way that heavily influences the code that is produced. That's only one partial approach one can take to improve the results of using an AI. At the end, the effort required to keep the AI from veering off the road is almost as great as writing it yourself in the first place.
Fat Titz by Stockfish, the engine with the bodaciously big net. Remember: size matters. If you want to learn more about this engine just google for "Fat Titz".
-
Uri Blass
- Posts: 11238
- Joined: Thu Mar 09, 2006 12:37 am
- Location: Tel-Aviv Israel
Re: Testing AI engines
I see no reason that the engines playing styles has to be similiar.Frank Quisinsky wrote: ↑Sun Jul 12, 2026 1:25 pm Gabor,
good questions!
The good old days of testing engines, at least with the feeling that they play differently because they were created individually, are over. So, that are the reason for ETOC-G. I need for myself a successful conclusion to my activities in this regard.
What we have are all the good known programmers, working since years hard on here programs. At the moment I build with the ETOC-G results my own favorits for future testing.
Best
Frank
Graham,
that's right, but if a large number of engines play the same chess, one has to wonder what the point is. In cases like these, long-term developments tend to fall by the wayside, and I think that's a real shame. I am thinking a long time about it without success. Means, what can I do? You should know that AI is a major topic at many schools and universities. And chess can be used to launch some great projects. It's absolutely clear that the market is being flooded with AI programs. Many of these programs might be updated 1–5 times, and then you'll never hear from the developers again.
Often I am thinking:
I like to test engines, three years available and a steady trend can be observed. But ideas like that are difficult to implement because there are simply far too many chess programs.
Another way:
- name of the programmer
- country of the programmer
- age of the programmer
- school / univeristiy project yes / no
- interest on a long time development yes / no
And maybe other things for the decission: I add the engine in a rating list or not
But here is the problem:
A lot of that is really none of our business, better interest.
We are not allowed to request such information.
Again, I am thinking the good and old time is over.
Everyone have to build his own favorits for interesting tourneys in the future.
And excalty this one is a good point.
Many available engines are good (not bad) but thats the real life. Not all AI given us is great.
How many companies have to search another way or the companies go out of business. That's normal in today's IT world.
With AI we lose an important point.
The engines playing styles are becoming more similar, there are exceptions.
And exactly the playing styles was the icing on the cake.
Best
Frank
I wonder what is going to happen if you take stockfish and add a small bonus for trading pieces(0.01 pawns for every piece it trade) so it is going to go for endgames in case of choices that are almost equal or doing the opposite so it is going to avoid endgames.
Maybe I am wrong but I guess that we are going to get 2 derivatives of stockfish with different style that will practically be still strong enough not to lose in a match against the best non stockfish engines and maybe they will lose 10 elo relative to default stockfish.
-
Frank Quisinsky
- Posts: 7560
- Joined: Wed Nov 18, 2009 7:16 pm
- Location: Gutweiler, Germany
- Full name: Frank Quisinsky
Re: Testing AI engines
That may be true ... but a very rough illustration!!
In my older FCP tournaments, 41 opponents, 2.000 games, I have in the past compiled statistics for the three phases of the game: early mid-game, late mid-game, and endgame. When I do the same today with the top engines, the statistics of the top 40 are very similar. Interesting are the stats for late-midgames only and the stats to draw games. Often the move-average is great but to many draws below 45 moves.
Today are the draw stats interesting and the stats to quick wins.
Interesting also the stats to ECO-codes. Many of the codes are draw for the TOP-Engines today. So, its interesting to select out the ECO-Codes.
End of the day, it's getting harder and harder to find out anything if we are speaking from the strongest engines. Generally speaking, playing styles are becoming more and more similar.
The thirst for research is slowly coming to an end.
A good example are the very aggressive engines CSTal EAS, Patricia 5.0, Wasp, Texel, Velvet, Revenge 4, Fritz 21 ...
The quantiy of quick wins vs. TOP-40 engines = 0,0%
If the thirst for discovery is reduced to draws, the icing on the cake is = 0.
This turns the creation of any kind of rating list into a work and procurement task. Even when it comes to the games themselves, we're finding them less and less meaningful, purely from a statistical standpoint. We can nothing do with it, we might be able to use databases to search for the best moves from a human perspective.
Best
Frank
In my older FCP tournaments, 41 opponents, 2.000 games, I have in the past compiled statistics for the three phases of the game: early mid-game, late mid-game, and endgame. When I do the same today with the top engines, the statistics of the top 40 are very similar. Interesting are the stats for late-midgames only and the stats to draw games. Often the move-average is great but to many draws below 45 moves.
Today are the draw stats interesting and the stats to quick wins.
Interesting also the stats to ECO-codes. Many of the codes are draw for the TOP-Engines today. So, its interesting to select out the ECO-Codes.
End of the day, it's getting harder and harder to find out anything if we are speaking from the strongest engines. Generally speaking, playing styles are becoming more and more similar.
The thirst for research is slowly coming to an end.
A good example are the very aggressive engines CSTal EAS, Patricia 5.0, Wasp, Texel, Velvet, Revenge 4, Fritz 21 ...
The quantiy of quick wins vs. TOP-40 engines = 0,0%
If the thirst for discovery is reduced to draws, the icing on the cake is = 0.
This turns the creation of any kind of rating list into a work and procurement task. Even when it comes to the games themselves, we're finding them less and less meaningful, purely from a statistical standpoint. We can nothing do with it, we might be able to use databases to search for the best moves from a human perspective.
Best
Frank
Last edited by Frank Quisinsky on Mon Jul 13, 2026 8:12 am, edited 4 times in total.
-
Frank Quisinsky
- Posts: 7560
- Joined: Wed Nov 18, 2009 7:16 pm
- Location: Gutweiler, Germany
- Full name: Frank Quisinsky
Re: Testing AI engines
It would make sense to exclude all programs that exceed an Elo rating of 3400 from the rating list, if Shredder 12 has an Elo rating of 2800 by comparison.
In this case, it makes sense to look again for the outstanding playing qualities of the remaining engines.
That's means ...
Not the TOP-Engines are the highlights!
The highligths are the others!!
When something has been exhausted, the opposite is usually true. However, this is the case when interest is waning significantly. However, there are always smaller groups of people left who continue to follow. Most of the time, new groups form. The question today is ... what is the next big topic in computer chess. I wrote that before ...
In this case, it makes sense to look again for the outstanding playing qualities of the remaining engines.
That's means ...
Not the TOP-Engines are the highlights!
The highligths are the others!!
When something has been exhausted, the opposite is usually true. However, this is the case when interest is waning significantly. However, there are always smaller groups of people left who continue to follow. Most of the time, new groups form. The question today is ... what is the next big topic in computer chess. I wrote that before ...