Fable 5.1: 3277 elo UCI NNUE Engine in 24 hrs...
Moderator: Ras
-
Sylwy
- Posts: 5373
- Joined: Fri Apr 21, 2006 4:19 pm
- Location: IAȘI - the historical capital of MOLDOVA
- Full name: Silvian Rucsandescu
-
op12no2
- Posts: 572
- Joined: Tue Feb 04, 2014 12:25 pm
- Location: Gower, Wales
- Full name: Colin Jenkins
Re: Fable 5.1: 3277 elo UCI NNUE Engine in 24 hrs...
I agree that it'd be interesting to see any references, given you allowed it to search. Also interesting would be a repeat experiment with explicit instructions not to search - just work from it's "memory".
-
Peter Berger
- Posts: 844
- Joined: Thu Mar 09, 2006 2:56 pm
Re: Fable 5.1: 3277 elo UCI NNUE Engine in 24 hrs...
I admire your prompt. It is easy to grasp how elegant it is.
I had an LLM explain Fable's approach to me in terms I could understand, and it pointed out two things that I found particularly interesting:
1. The time Fable spent trying to understand the strange observation that giving it more time per move actually made it play worse, eventually tracing this to the 16-bit hash verification issue.
2. The way it handled the risk involved in developing and eventually adopting the NNUE. This one looked amazingly intelligent to me.
I had an LLM explain Fable's approach to me in terms I could understand, and it pointed out two things that I found particularly interesting:
1. The time Fable spent trying to understand the strange observation that giving it more time per move actually made it play worse, eventually tracing this to the 16-bit hash verification issue.
2. The way it handled the risk involved in developing and eventually adopting the NNUE. This one looked amazingly intelligent to me.
-
chessica
- Posts: 1140
- Joined: Thu Aug 11, 2022 11:30 pm
- Full name: Esmeralda Pinto
Re: Fable 5.1: 3277 elo UCI NNUE Engine in 24 hrs...
Code: Select all
Program Elo + - Games Score Av.Op. Draws
1 Stockfish 18 : 3555 22 21 573 80.0 % 3314 40.0 %
2 Revolution-5.90-140626 : 3553 22 21 576 79.9 % 3314 39.9 %
3 Cool Iris 16 : 3552 20 19 700 79.3 % 3319 41.4 %
4 Reckless 0.9.0 : 3551 22 21 573 79.7 % 3314 40.7 %
5 Reckless 0.10.0-dev : 3543 25 24 470 80.5 % 3297 38.9 %
6 Rems M-091024 : 3534 21 20 572 78.0 % 3314 43.7 %
7 Obsidian dev-16.08 : 3529 22 21 575 77.5 % 3315 43.0 %
8 Coda 0.9.4 : 3529 25 24 470 79.1 % 3297 40.0 %
9 PlentyChess 7.0.26 : 3525 19 19 700 76.6 % 3320 44.0 %
10 PlentyChess 8.0.0 : 3521 24 23 470 78.4 % 3297 42.8 %
11 Tarnished v6.0 (Eternal) : 3515 22 21 573 76.0 % 3315 43.1 %
12 Berserk 14 : 3510 21 21 576 75.4 % 3315 44.3 %
13 Viridithas 20.0.0-dev : 3486 22 22 572 72.8 % 3315 41.4 %
14 RubiChess 20240817 (avx512) : 3480 21 20 645 71.9 % 3317 42.6 %
15 Igel 3.7.0 64 POPCNT AVX2 : 3470 24 23 470 72.9 % 3298 43.6 %
16 Rebel 16.3a : 3459 20 20 701 68.8 % 3322 40.4 %
17 Chess System Tal 2.06 E1019 : 3451 22 22 575 68.4 % 3317 41.2 %
18 RubiChess 20240817 (avx2) : 3447 20 20 685 67.2 % 3322 41.8 %
19 Maverick NNUE 1.0a : 3404 20 20 702 61.5 % 3323 39.2 %
20 Alexander 7 : 3400 23 23 572 61.7 % 3317 37.8 %
21 Shredder Classic 6 : 3378 26 26 471 61.0 % 3300 32.9 %
22 Willow 4.0 : 3368 23 23 572 57.3 % 3317 33.0 %
23 Winter 4.0 JA BMI2 : 3358 21 21 704 54.8 % 3324 31.2 %
24 Cheng 4.48a (bundled) : 3339 24 24 575 53.0 % 3319 30.1 %
25 Houdidit 6.03 Pro x64 : 3328 27 27 470 53.8 % 3301 26.8 %
26 Komodo 14.1 64-bit : 3326 24 24 575 51.0 % 3319 30.3 %
27 EveAnn 5.0 64-bit (bundled) : 3270 25 25 573 42.9 % 3320 25.1 %
28 Myrddin 0.96 : 3253 25 25 573 40.5 % 3320 22.7 %
29 Fable 5.1 : 3235 29 29 470 40.3 % 3303 16.0 %
30 IsaBB NN 4.5 (avx2, bundled) : 3233 26 26 572 37.7 % 3320 22.6 %
31 Sable 1.6 : 3205 27 27 573 33.9 % 3321 16.2 %
32 Equinox 3.30 x64mp : 3199 29 29 470 35.3 % 3304 22.1 %
33 Critter 1.6a 64-bit : 3190 27 28 572 31.9 % 3321 16.3 %
34 Houdini 1.5a x64 : 3173 31 31 474 31.6 % 3306 13.1 %
35 OpenCritter v1.1.37 : 3162 30 30 470 30.5 % 3305 19.8 %
36 ECE 26.8 Intuition HBuP : 3141 32 32 470 28.0 % 3305 13.0 %
37 monty 1.0.0 : 3136 30 30 575 25.4 % 3323 12.2 %
38 Crafty 25.2.1 : 3100 28 28 702 20.9 % 3331 14.0 %
39 Spike 1.4.2 : 3094 33 34 470 22.8 % 3306 13.6 %
40 Spike 1.4.1 : 3092 31 32 572 20.9 % 3323 12.8 %
41 Colossus Chess 2026a : 3049 37 38 470 18.4 % 3307 9.6 %
42 Rybka 1.0 Beta : 3049 36 37 470 18.4 % 3307 10.9 %
43 Colossus 2025b : 3036 34 35 572 16.0 % 3324 11.7 %
44 Pharaon 3.5.1 : 2883 49 51 573 7.2 % 3328 4.9 %
45 SOS 5 for Arena : 2836 60 63 470 6.1 % 3312 3.2 %
46 Quark v2.35Paderborn : 2805 65 69 470 5.1 % 3312 2.6 %
47 Yace 0.99.87 : 2793 65 70 470 4.8 % 3313 3.2 %
-
chrisw
- Posts: 5112
- Joined: Tue Apr 03, 2012 4:28 pm
- Location: Digital Nomad. Anywhere but the Western Empire
- Full name: Christopher Whittington
Re: Fable 5.1: 3277 elo UCI NNUE Engine in 24 hrs...
Impressive! Did the models run using their own resources? Including building games databases?Steve Maughan wrote: ↑Thu Sep 03, 2026 3:50 am I've been thinking about a coding benchmark for large language models, and a chess engine is close to ideal for it. The result is objective, it is measurable, and you can put the engines on a board against each other and see who wins. The test is simple. The model gets 24 hours of wall-clock time, the UCI spec, fastchess with the UHO book, a handful of Stash builds as a ladder and a perft suite, and it must deliver a UCI engine from scratch. Public information is fine (chessprogramming.org, papers, PeSTO tables), but no code from existing engines, no pre-existing nets or books, standard library only, C or C++. Nobody answers questions during the run. The only thing scored is Elo at 10s+0.1s. The idea was partly inspired by the CODA project, which took months and produced a very strong engine; I wanted to see what a model could do in a single day.
I ran four Claude models. Each engine was then rated in one 3,800-game run against Stash 20 to 37, Crafty 25.6 and Juggernaut, anchored to CCRL Blitz:
(CCRL Blitz scale under 10+0.1 conditions: one thread, 64 MB hash, UHO openings played with both colours.)Code: Select all
Fable 5.1 3277 ±23 Opus 5 3242 ±23 Fable 5 3049 ±22 Sonnet 5 2702 ±26
Opus 5 producing a 3200+ engine in a day surprised me. Fable 5 did less well than I expected, and Sonnet 5 landed around 2700. But Fable 5.1, which I ran only yesterday, is the one that blew me away. I had assumed 24 hours left room for nothing more than a hand-crafted (well, AI-crafted) evaluation. About five and a half hours in, Fable 5.1 wrote its own self-play data generator and an NNUE trainer. Nothing of the sort was provided. By hour ten the net had reached parity with its hand-crafted eval, and the final engine ships a 768→384x2 network trained on roughly 17 million of its own self-play positions, weights compiled into the executable, with the hand-crafted eval left in as a fallback. 3277 Elo.
The other thing that amazed me was sheer speed. Each model logged when it first passed the full perft suite (126 positions, depths 1 to 6) and when its engine first played a complete game. Sonnet 5 had a bitboard, PEXT move generator passing every perft position 6 minutes after the clock started, on the first run, with no bugs to fix. Fable 5.1 passed perft at 8 minutes and by 15 minutes had a compliance-checked build in final/ playing a 10+0.1 match against Stash 20. Fable 5 was a minute or two behind on both. I have spent longer than that looking for a castling bug.
The repositories are below. Each has a README the model wrote after the deadline, the hourly progress log it kept during the run, and a release with the executable.
I expect some here won't like this. It adds to the flood of engines, and I have some sympathy with that view. But the flood is already here; look at the entry lists for recent CCRL tournaments. In the summer of 2025, while I was working on Juggernaut, I asked Graham Banks what it took to get into his tournaments, and the answer was about 2450 for the bottom division. Juggernaut is now 2750 and only just strong enough. That is not why I built this, though. I built it because it is a genuine coding benchmark, objective and measurable, and the bar will only rise. Perhaps we will see a Stockfish written in 24 hours, which would be both amazing and a little terrible. I feel both. Comments and input welcome.
- Fable 5.1: https://github.com/stevemaughan/fable51-chess-24hrs
- Opus 5: https://github.com/stevemaughan/opus5-chess-24hrs
- Fable 5: https://github.com/stevemaughan/fable5-chess-24hrs
- Sonnet 5: https://github.com/stevemaughan/sonnet5-chess-24hrs
- The benchmark itself, for anyone who wants to run another model: https://github.com/stevemaughan/chess-engine-benchmark
Steve
-
Steve Maughan
- Posts: 1365
- Joined: Wed Mar 08, 2006 8:28 pm
- Location: Florida, USA
Re: Fable 5.1: 3277 elo UCI NNUE Engine in 24 hrs...
Yes! Only Fable 5.1 attempted to create a NNUE evaluation function. At around 5 hours it started generating positions and coding a NNUE training routine. I thought NNUE wasn't possible in a 24 hour timeframe — I was wrong. The other interesting aspect of Fable 5.1's run is the way it managed continuously improved. It's first engine was only 2200 elo but it steadily climbed up the ratings. The other AIs didn't improve as much after they had a working engine.
— Steve
http://www.chessprogramming.net - Juggernaut & Maverick Chess Engine
-
glav
- Posts: 97
- Joined: Sun Apr 07, 2019 1:10 am
- Full name: Giovanni Lavorgna
Re: Fable 5.1: 3277 elo UCI NNUE Engine in 24 hrs...
Interesting. Which hardware were you using? Roughly how many games you playes to test each Elo increase?Steve Maughan wrote: ↑Fri Sep 04, 2026 10:20 pm
Yes! Only Fable 5.1 attempted to create a NNUE evaluation function. At around 5 hours it started generating positions and coding a NNUE training routine. I thought NNUE wasn't possible in a 24 hour timeframe — I was wrong. The other interesting aspect of Fable 5.1's run is the way it managed continuously improved. It's first engine was only 2200 elo but it steadily climbed up the ratings. The other AIs didn't improve as much after they had a working engine.
— Steve
-
brianr
- Posts: 542
- Joined: Thu Mar 09, 2006 3:01 pm
- Full name: Brian Richardson
Re: Fable 5.1: 3277 elo UCI NNUE Engine in 24 hrs...
Thank you for posting.
This is amazing.
Can you provide some $cost or token cost information?
This is amazing.
Can you provide some $cost or token cost information?
-
Steve Maughan
- Posts: 1365
- Joined: Wed Mar 08, 2006 8:28 pm
- Location: Florida, USA
Re: Fable 5.1: 3277 elo UCI NNUE Engine in 24 hrs...
The test was carried out on a three year old 13th generation Intel i7 laptop with 10 cores. This is relevant since it determines the concurrency of the tests carried out by the AI.
Note: I didn’t do any testing — it was all done by the AI. This is a single prompt test.
— Steve
http://www.chessprogramming.net - Juggernaut & Maverick Chess Engine
-
Steve Maughan
- Posts: 1365
- Joined: Wed Mar 08, 2006 8:28 pm
- Location: Florida, USA
Re: Fable 5.1: 3277 elo UCI NNUE Engine in 24 hrs...
I’m on the $200 max plan. At no time was it close to maxing out in any five hour period — maybe 20%. The truth is, the AI isn’t coding that much of the time. Most of the time is needed to test changes.
— Steve
http://www.chessprogramming.net - Juggernaut & Maverick Chess Engine