← All experiments THE RESEARCH RECORD / cudax-growth-heldout-rerun

CUDAX progressive-growth held-out rerun

Under the YAML harness with a tail held-out split, does giving the 16-division progressive-growth schedule more epochs per frontier improve held-out byte PPL?

experimentconcludedmixedattentioneval: canonical

Run history: 3 of 3 predate the 2026-09-24 CUDA kernel race fix; 3 of 3 predate the 2026-09-28 exact ancestor backward.

Opened Updated
THE ANSWER SO FAR

Up to 3 epochs per frontier. Going from 1 to 3 improved rolling byte PPL from 9.57 to 8.56, and 6 epochs gave no further gain (8.57).

progressive 16 x 3 epochs 8.5556 rolling byte PPL (tail-heldout) Run record ↗pre-race-fixtruncated-ancestor-gradient
progressive 16 x 6 epochs 8.5672 rolling byte PPL (tail-heldout) Run record ↗pre-race-fixtruncated-ancestor-gradient
progressive 16 x 1 epoch 9.5708 rolling byte PPL (tail-heldout) Run record ↗pre-race-fixtruncated-ancestor-gradient
TOPICS
RELATED
SUPERSEDED BY
CODE
THE FULL RECORD

Experiment notes

Open directory on GitHub ↗

cudax-growth-heldout-rerun

Status: concluded, mixed (reviewed 2026-09-28; this line previously read "active"). The answer is in the front matter above.

Hypothesis

The earlier CUDAX progressive-growth numbers should be reproducible under the YAML experiment harness with an explicit held-out tail split. If the progressive schedule is useful, increasing the number of epochs per growth frontier should improve held-out byte PPL, but gains may saturate quickly.

Scope

  • Trainer: bin/agpt_train_v2
  • Mode: train-growth
  • Growth divisions: 16
  • Epochs per frontier: 1, 3, 6
  • Corpus: data/input.txt
  • Training split: prefix 95%
  • Evaluation split: held-out tail 5%
  • Model init: data/input.model
  • Depth/window: growth_max_depth: 16
  • Optimizer: RMSProp, lr=0.003, rmsprop_beta=0.999
  • Schedule: warmup-cosine, warmup_epochs=0
  • Ancestor gradients: enabled

Results

<!-- agpt-experiment-table:start -->

Run ID byte_ppl bits/byte word_ppl wall (s)
20260526T041157-progressive-16x1 9.5708 3.2586 305118.35 85.0
20260526T041333-progressive-16x3 8.5556 3.0969 162996.4 183.0
20260526T041703-progressive-16x6 8.5672 3.0988 164240.94 394.0
<!-- agpt-experiment-table:end -->

Conclusion

The 3-epoch and 6-epoch progressive runs land close together on held-out byte PPL: 8.5556 vs 8.5672. The 1-epoch run is worse at 9.5708. In this batch, extra work beyond 3 epochs per frontier did not produce a clear held-out improvement.