← All experiments THE RESEARCH RECORD / cudax-d16-linear-mass-rerun

CUDAX d16 linear-mass rerun

What tail-heldout PPL does the v2 trainer's d=16 trie, d64/L2 static-epoch baseline reach with paper-correct linear mass weighting at 25, 100 and 250 epochs?

baselineconcludedn/aattentioneval: canonical

Run history: 5 of 5 predate the 2026-09-24 CUDA kernel race fix; 5 of 5 predate the 2026-09-28 exact ancestor backward.

Opened Updated
THE ANSWER SO FAR

100 epochs was best (rolling byte PPL 6.98). 250 epochs was slightly worse (7.15) and 25 epochs worse (8.44). Two later reruns of the 100-epoch config under the migrated config schema gave 7.27 and 7.12. No conclusion was written.

static 100 epochs 6.9773 rolling byte PPL (tail-heldout) Run record ↗pre-race-fixtruncated-ancestor-gradient
static 250 epochs 7.1471 rolling byte PPL (tail-heldout) Run record ↗pre-race-fixtruncated-ancestor-gradient
static 25 epochs 8.439 rolling byte PPL (tail-heldout) Run record ↗pre-race-fixtruncated-ancestor-gradient
TOPICS
RELATED
SUPERSEDED BY
CODE
THE FULL RECORD

Experiment notes

Open directory on GitHub ↗

cudax-d16-linear-mass-rerun

Status: concluded, n/a (reviewed 2026-09-28; this line previously read "active"). The answer is in the front matter above.

Hypothesis

(fill in)

Scope

(fill in)

Results

<!-- agpt-experiment-table:start -->

Run ID byte_perplexity bits/byte train (s) total (s)
20260528T165653-d16-d64l2-static25 — — 152.0 185.0
20260528T170108-d16-d64l2-static100 — — 634.0 663.0
20260528T171523-d16-d64l2-static250 — — 1510.0 1543.0
20260528T230843-d16-d64l2-static100 — — 552.0 580.0
20260529T001452-d16-d64l2-static100 — — 599.0 628.0
<!-- agpt-experiment-table:end -->

Conclusion

(fill in once enough runs have landed)