This calculation effort ran mostly on a single 8x 5090 machine, with 512gb of host memory
Raw GPU time (in terms of 1x gpu hours, including completely idle time, test runs, iterations)
Code: Select all
RTX 5090 x8: 203h 20m 9s
RTX 5090 x4: 182h 15m 41s
RTX 5090: 3h 16m 31sStart - End
Code: Select all
[2026-08-16 13:18:48:453] - [GPU] 63.7% 168495518731/~264432813441 roots 14.0s elapsed 14687155.5 GN/s 50046.37 raw GN/s pool 347/378 ~ETA 40739.2s (60s 147546 raw)
[2026-08-17 00:49:36:442] - [GPU] 100.0% 208523003998/~208523003998 roots 41462.0s elapsed 13615389.4 GN/s 58960.15 raw GN/s pool 0/1 ETA -- (60s 51348 raw)Interesting in that effort is the dedup tree, which is never fully materialized to fit withing constrained memory.
Code: Select all
[2026-08-17 00:49:52:220] - dedup tree (host roots):
[2026-08-17 00:49:52:220] - |- ply 1: 20 unique 20 nodes 1.00x dedup
[2026-08-17 00:49:52:220] - |- ply 2: 400 unique 400 nodes 1.00x dedup
[2026-08-17 00:49:52:220] - |- ply 3: 5362 unique 8902 nodes 1.66x dedup
[2026-08-17 00:49:52:220] - |- ply 4: 72078 unique 197281 nodes 2.74x dedup
[2026-08-17 00:49:52:220] - |- ply 5: 822518 unique 4865609 nodes 5.92x dedup
[2026-08-17 00:49:52:220] - |- ply 6: 9417683 unique 119060324 nodes 12.64x dedup
[2026-08-17 00:49:52:220] - |- ply 7: 96400335 unique 3195901860 nodes 33.15x dedup
[2026-08-17 00:49:52:220] - |- ply 8: 988192872 unique 84998978956 nodes 86.01x dedup
[2026-08-17 00:49:52:220] - |- ply 9: fused 2439530234167 nodes (streamed without materializing)
[2026-08-17 00:49:52:220] - \- ply 10: 75670835200 staged 69352859712417 nodes 916.51x dedup (streamed)viewtopic.php?t=86588
This is the first independent verification of 2,015,099,950,053,364,471,960 and I can confirm correctness.
Signed - Daniel Infuehr
Also - P16 is totally reachable by more than 20x this effort. 20x+ since we will start one ply deep.
If you have access to a gpu cluster PM me.