← All experiments THE RESEARCH RECORD / window-d124-baseline

Window baseline at seq_len 124

What held-out PPL does a standard sliding-window microgpt model reach at seq_len 124 (d64 L2, about 25 corpus epochs), as a reference for depth-124 AGPT?

baselineconcludedn/aattentioneval: canonical

Run history: 1 of 1 predate the 2026-09-24 CUDA kernel race fix; 1 of 1 predate the 2026-09-28 exact ancestor backward.

Opened Updated
THE ANSWER SO FAR

Rolling byte PPL 5.915 and fixed-token PPL 5.838 on the 5% tail heldout after 214,000 steps at constant lr 3e-4.

microgpt window d64/L2 seq 124, 214k steps 5.915 rolling byte PPL (tail-heldout 5%) Run record ↗pre-race-fixtruncated-ancestor-gradient
microgpt window d64/L2 seq 124, 214k steps 5.8381 fixed-window PPL (tail-heldout 5%) Run record ↗pre-race-fixtruncated-ancestor-gradient
TOPICS
RELATED
SUPERSEDED BY
CODE
THE FULL RECORD

Experiment notes

Open directory on GitHub ↗

window-d124-baseline

Status: concluded, n/a (reviewed 2026-09-28; this line previously read "active"). The answer is in the front matter above.

Hypothesis

(fill in)

Scope

(fill in)

Results

<!-- agpt-experiment-table:start -->

Run ID fixed_token_ppl rolling_byte_ppl bits/byte train (s) total (s)
20260528T094311-window-adam-d64l2-s124-25ep 5.8381 5.915 2.5644 4292.0 4330.0
<!-- agpt-experiment-table:end -->

Conclusion

(fill in once enough runs have landed)