← All experiments THE RESEARCH RECORD / shake-sgd-baseline

Shakespeare SGD window baseline

What held-out PPL does a standard sliding-window microgpt model (d64 L2, seq_len 16, 10k steps) reach under the AGPT experiment-harness split and evaluator?

baselineconcludedn/aattentioneval: canonical

Run history: 1 of 1 predate the 2026-09-24 CUDA kernel race fix; 1 of 1 predate the 2026-09-28 exact ancestor backward.

Opened Updated
THE ANSWER SO FAR

Rolling byte PPL 10.23 and fixed-token PPL 10.06 on the 5% tail heldout after 10,000 steps at constant lr 3e-4.

microgpt window d64/L2 seq 16, 10k steps 10.2344 rolling byte PPL (tail-heldout 5%) Run record ↗pre-race-fixtruncated-ancestor-gradient
microgpt window d64/L2 seq 16, 10k steps 10.0631 fixed-window PPL (tail-heldout 5%) Run record ↗pre-race-fixtruncated-ancestor-gradient
TOPICS
RELATED
SUPERSEDED BY
CODE
THE FULL RECORD

Experiment notes

Open directory on GitHub ↗

shake-sgd-baseline

Status: initial baseline landed

Question

Record a standard Crystal microgpt sliding-window baseline for the Shakespeare small-model setup, using the same AGPT experiment harness split and evaluator.

Protocol

  • Trainer: /home/trans/Projects/microgpt/bin/microgpt
  • Mode: standard sliding-window SGD
  • Corpus: data/input.txt
  • Training split: prefix 95%
  • Evaluation split: held-out tail 5%
  • Model init: data/input.model
  • Window: seq_len=16
  • Model: d_model=64, n_layers=2, d_ff=256
  • Optimizer: SGD, lr=0.0003, constant schedule
  • Budget: 10,000 steps

Caveats

This run predates the stricter raw-log policy and did not originally write a result.json; the committed result.json was reconstructed from eval_raw.json and meta.json.

Results

Run ID fixed_token_ppl rolling_byte_ppl bits/byte
20260528T181839-d64l2-s16-10k 10.0631 10.2344 3.3553