Fable 5.1: 3277 elo UCI NNUE Engine in 24 hrs...

Discussion of anything and everything relating to chess playing software and machines.

Moderator: Ras

User avatar
Sylwy
Posts: 5373
Joined: Fri Apr 21, 2006 4:19 pm
Location: IAȘI - the historical capital of MOLDOVA
Full name: Silvian Rucsandescu

Re: Fable 5.1: 3277 elo UCI NNUE Engine in 24 hrs...

Post by Sylwy »

op12no2
Posts: 572
Joined: Tue Feb 04, 2014 12:25 pm
Location: Gower, Wales
Full name: Colin Jenkins

Re: Fable 5.1: 3277 elo UCI NNUE Engine in 24 hrs...

Post by op12no2 »

I agree that it'd be interesting to see any references, given you allowed it to search. Also interesting would be a repeat experiment with explicit instructions not to search - just work from it's "memory".
Peter Berger
Posts: 844
Joined: Thu Mar 09, 2006 2:56 pm

Re: Fable 5.1: 3277 elo UCI NNUE Engine in 24 hrs...

Post by Peter Berger »

I admire your prompt. It is easy to grasp how elegant it is.

I had an LLM explain Fable's approach to me in terms I could understand, and it pointed out two things that I found particularly interesting:

1. The time Fable spent trying to understand the strange observation that giving it more time per move actually made it play worse, eventually tracing this to the 16-bit hash verification issue.
2. The way it handled the risk involved in developing and eventually adopting the NNUE. This one looked amazingly intelligent to me.
chessica
Posts: 1140
Joined: Thu Aug 11, 2022 11:30 pm
Full name: Esmeralda Pinto

Re: Fable 5.1: 3277 elo UCI NNUE Engine in 24 hrs...

Post by chessica »

Code: Select all


    Program                          Elo    +   -   Games   Score   Av.Op.  Draws

  1 Stockfish 18                   : 3555   22  21   573    80.0 %   3314   40.0 %
  2 Revolution-5.90-140626         : 3553   22  21   576    79.9 %   3314   39.9 %
  3 Cool Iris 16                   : 3552   20  19   700    79.3 %   3319   41.4 %
  4 Reckless 0.9.0                 : 3551   22  21   573    79.7 %   3314   40.7 %
  5 Reckless 0.10.0-dev            : 3543   25  24   470    80.5 %   3297   38.9 %
  6 Rems M-091024                  : 3534   21  20   572    78.0 %   3314   43.7 %
  7 Obsidian dev-16.08             : 3529   22  21   575    77.5 %   3315   43.0 %
  8 Coda 0.9.4                     : 3529   25  24   470    79.1 %   3297   40.0 %
  9 PlentyChess 7.0.26             : 3525   19  19   700    76.6 %   3320   44.0 %
 10 PlentyChess 8.0.0              : 3521   24  23   470    78.4 %   3297   42.8 %
 11 Tarnished v6.0 (Eternal)       : 3515   22  21   573    76.0 %   3315   43.1 %
 12 Berserk 14                     : 3510   21  21   576    75.4 %   3315   44.3 %
 13 Viridithas 20.0.0-dev          : 3486   22  22   572    72.8 %   3315   41.4 %
 14 RubiChess 20240817 (avx512)    : 3480   21  20   645    71.9 %   3317   42.6 %
 15 Igel 3.7.0 64 POPCNT AVX2      : 3470   24  23   470    72.9 %   3298   43.6 %
 16 Rebel 16.3a                    : 3459   20  20   701    68.8 %   3322   40.4 %
 17 Chess System Tal 2.06 E1019    : 3451   22  22   575    68.4 %   3317   41.2 %
 18 RubiChess 20240817 (avx2)      : 3447   20  20   685    67.2 %   3322   41.8 %
 19 Maverick NNUE 1.0a             : 3404   20  20   702    61.5 %   3323   39.2 %
 20 Alexander 7                    : 3400   23  23   572    61.7 %   3317   37.8 %
 21 Shredder Classic 6             : 3378   26  26   471    61.0 %   3300   32.9 %
 22 Willow 4.0                     : 3368   23  23   572    57.3 %   3317   33.0 %
 23 Winter 4.0 JA BMI2             : 3358   21  21   704    54.8 %   3324   31.2 %
 24 Cheng 4.48a (bundled)          : 3339   24  24   575    53.0 %   3319   30.1 %
 25 Houdidit 6.03 Pro x64          : 3328   27  27   470    53.8 %   3301   26.8 %
 26 Komodo 14.1 64-bit             : 3326   24  24   575    51.0 %   3319   30.3 %
 27 EveAnn 5.0 64-bit (bundled)    : 3270   25  25   573    42.9 %   3320   25.1 %
 28 Myrddin 0.96                   : 3253   25  25   573    40.5 %   3320   22.7 %
 29 Fable 5.1                      : 3235   29  29   470    40.3 %   3303   16.0 %
 30 IsaBB NN 4.5 (avx2, bundled)   : 3233   26  26   572    37.7 %   3320   22.6 %
 31 Sable 1.6                      : 3205   27  27   573    33.9 %   3321   16.2 %
 32 Equinox 3.30 x64mp             : 3199   29  29   470    35.3 %   3304   22.1 %
 33 Critter 1.6a 64-bit            : 3190   27  28   572    31.9 %   3321   16.3 %
 34 Houdini 1.5a x64               : 3173   31  31   474    31.6 %   3306   13.1 %
 35 OpenCritter v1.1.37            : 3162   30  30   470    30.5 %   3305   19.8 %
 36 ECE 26.8 Intuition HBuP        : 3141   32  32   470    28.0 %   3305   13.0 %
 37 monty 1.0.0                    : 3136   30  30   575    25.4 %   3323   12.2 %
 38 Crafty 25.2.1                  : 3100   28  28   702    20.9 %   3331   14.0 %
 39 Spike 1.4.2                    : 3094   33  34   470    22.8 %   3306   13.6 %
 40 Spike 1.4.1                    : 3092   31  32   572    20.9 %   3323   12.8 %
 41 Colossus Chess 2026a           : 3049   37  38   470    18.4 %   3307    9.6 %
 42 Rybka 1.0 Beta                 : 3049   36  37   470    18.4 %   3307   10.9 %
 43 Colossus 2025b                 : 3036   34  35   572    16.0 %   3324   11.7 %
 44 Pharaon 3.5.1                  : 2883   49  51   573     7.2 %   3328    4.9 %
 45 SOS 5 for Arena                : 2836   60  63   470     6.1 %   3312    3.2 %
 46 Quark v2.35Paderborn           : 2805   65  69   470     5.1 %   3312    2.6 %
 47 Yace 0.99.87                   : 2793   65  70   470     4.8 %   3313    3.2 %
User avatar
chrisw
Posts: 5112
Joined: Tue Apr 03, 2012 4:28 pm
Location: Digital Nomad. Anywhere but the Western Empire
Full name: Christopher Whittington

Re: Fable 5.1: 3277 elo UCI NNUE Engine in 24 hrs...

Post by chrisw »

Steve Maughan wrote: Thu Sep 03, 2026 3:50 am I've been thinking about a coding benchmark for large language models, and a chess engine is close to ideal for it. The result is objective, it is measurable, and you can put the engines on a board against each other and see who wins. The test is simple. The model gets 24 hours of wall-clock time, the UCI spec, fastchess with the UHO book, a handful of Stash builds as a ladder and a perft suite, and it must deliver a UCI engine from scratch. Public information is fine (chessprogramming.org, papers, PeSTO tables), but no code from existing engines, no pre-existing nets or books, standard library only, C or C++. Nobody answers questions during the run. The only thing scored is Elo at 10s+0.1s. The idea was partly inspired by the CODA project, which took months and produced a very strong engine; I wanted to see what a model could do in a single day.

I ran four Claude models. Each engine was then rated in one 3,800-game run against Stash 20 to 37, Crafty 25.6 and Juggernaut, anchored to CCRL Blitz:

Code: Select all

Fable 5.1   3277  ±23
Opus 5      3242  ±23
Fable 5     3049  ±22
Sonnet 5    2702  ±26
(CCRL Blitz scale under 10+0.1 conditions: one thread, 64 MB hash, UHO openings played with both colours.)

Opus 5 producing a 3200+ engine in a day surprised me. Fable 5 did less well than I expected, and Sonnet 5 landed around 2700. But Fable 5.1, which I ran only yesterday, is the one that blew me away. I had assumed 24 hours left room for nothing more than a hand-crafted (well, AI-crafted) evaluation. About five and a half hours in, Fable 5.1 wrote its own self-play data generator and an NNUE trainer. Nothing of the sort was provided. By hour ten the net had reached parity with its hand-crafted eval, and the final engine ships a 768→384x2 network trained on roughly 17 million of its own self-play positions, weights compiled into the executable, with the hand-crafted eval left in as a fallback. 3277 Elo.

The other thing that amazed me was sheer speed. Each model logged when it first passed the full perft suite (126 positions, depths 1 to 6) and when its engine first played a complete game. Sonnet 5 had a bitboard, PEXT move generator passing every perft position 6 minutes after the clock started, on the first run, with no bugs to fix. Fable 5.1 passed perft at 8 minutes and by 15 minutes had a compliance-checked build in final/ playing a 10+0.1 match against Stash 20. Fable 5 was a minute or two behind on both. I have spent longer than that looking for a castling bug.

The repositories are below. Each has a README the model wrote after the deadline, the hourly progress log it kept during the run, and a release with the executable.
I expect some here won't like this. It adds to the flood of engines, and I have some sympathy with that view. But the flood is already here; look at the entry lists for recent CCRL tournaments. In the summer of 2025, while I was working on Juggernaut, I asked Graham Banks what it took to get into his tournaments, and the answer was about 2450 for the bottom division. Juggernaut is now 2750 and only just strong enough. That is not why I built this, though. I built it because it is a genuine coding benchmark, objective and measurable, and the bar will only rise. Perhaps we will see a Stockfish written in 24 hours, which would be both amazing and a little terrible. I feel both. Comments and input welcome.

Steve
Impressive! Did the models run using their own resources? Including building games databases?
User avatar
Steve Maughan
Posts: 1365
Joined: Wed Mar 08, 2006 8:28 pm
Location: Florida, USA

Re: Fable 5.1: 3277 elo UCI NNUE Engine in 24 hrs...

Post by Steve Maughan »

chrisw wrote: Fri Sep 04, 2026 9:05 pm Impressive! Did the models run using their own resources? Including building games databases?
Yes! Only Fable 5.1 attempted to create a NNUE evaluation function. At around 5 hours it started generating positions and coding a NNUE training routine. I thought NNUE wasn't possible in a 24 hour timeframe — I was wrong. The other interesting aspect of Fable 5.1's run is the way it managed continuously improved. It's first engine was only 2200 elo but it steadily climbed up the ratings. The other AIs didn't improve as much after they had a working engine.

— Steve
http://www.chessprogramming.net - Juggernaut & Maverick Chess Engine
glav
Posts: 97
Joined: Sun Apr 07, 2019 1:10 am
Full name: Giovanni Lavorgna

Re: Fable 5.1: 3277 elo UCI NNUE Engine in 24 hrs...

Post by glav »

Steve Maughan wrote: Fri Sep 04, 2026 10:20 pm
Yes! Only Fable 5.1 attempted to create a NNUE evaluation function. At around 5 hours it started generating positions and coding a NNUE training routine. I thought NNUE wasn't possible in a 24 hour timeframe — I was wrong. The other interesting aspect of Fable 5.1's run is the way it managed continuously improved. It's first engine was only 2200 elo but it steadily climbed up the ratings. The other AIs didn't improve as much after they had a working engine.

— Steve
Interesting. Which hardware were you using? Roughly how many games you playes to test each Elo increase?
brianr
Posts: 542
Joined: Thu Mar 09, 2006 3:01 pm
Full name: Brian Richardson

Re: Fable 5.1: 3277 elo UCI NNUE Engine in 24 hrs...

Post by brianr »

Thank you for posting.
This is amazing.

Can you provide some $cost or token cost information?
User avatar
Steve Maughan
Posts: 1365
Joined: Wed Mar 08, 2006 8:28 pm
Location: Florida, USA

Re: Fable 5.1: 3277 elo UCI NNUE Engine in 24 hrs...

Post by Steve Maughan »

glav wrote: Fri Sep 04, 2026 10:30 pm Interesting. Which hardware were you using? Roughly how many games you playes to test each Elo increase?
The test was carried out on a three year old 13th generation Intel i7 laptop with 10 cores. This is relevant since it determines the concurrency of the tests carried out by the AI.

Note: I didn’t do any testing — it was all done by the AI. This is a single prompt test.

— Steve
http://www.chessprogramming.net - Juggernaut & Maverick Chess Engine
User avatar
Steve Maughan
Posts: 1365
Joined: Wed Mar 08, 2006 8:28 pm
Location: Florida, USA

Re: Fable 5.1: 3277 elo UCI NNUE Engine in 24 hrs...

Post by Steve Maughan »

brianr wrote: Fri Sep 04, 2026 11:27 pm Can you provide some $cost or token cost information?
I’m on the $200 max plan. At no time was it close to maxing out in any five hour period — maybe 20%. The truth is, the AI isn’t coding that much of the time. Most of the time is needed to test changes.

— Steve
http://www.chessprogramming.net - Juggernaut & Maverick Chess Engine