← All experiments THE RESEARCH RECORD / lr1e-5-probe

LR 1e-5 probe

How does 10-epoch RMSProp training on the Shakespeare d=16 radix trie compare at lr 1e-5 (with and without branching-endpoint entropy icing) against lr 3e-3 (constant and warmup-cosine)?

diagnosticconcludedn/aattentioneval: none
Opened Updated
THE ANSWER SO FAR

On training loss only: lr 1e-5 barely trains in 10 epochs (loss 3.393 with icing, 3.349 without), while lr 3e-3 reaches 2.199 with a constant schedule and 2.063 with warmup-cosine. No held-out PPL was measured.

lr 1e-5, entropy icing on, 10 epochs 3.393104 train loss (nats, training trie) pre-loss-fixpre-race-fixtruncated-ancestor-gradient
lr 3e-3 warmup-cosine, 10 epochs 2.063377 train loss (nats, training trie) pre-loss-fixpre-race-fixtruncated-ancestor-gradient
TOPICS
RELATED
SUPERSEDED BY
CODE
THE FULL RECORD

Experiment notes

Open directory on GitHub ↗

LR 1e-5 probe

Stub README (2026-09-28): this directory predates the README convention. The summary below was reconstructed from its files; see them for detail.

Question. How does 10-epoch RMSProp training on the Shakespeare d=16 radix trie compare at lr 1e-5 (with and without branching-endpoint entropy icing) against lr 3e-3 (constant and warmup-cosine)?

Answer. On training loss only: lr 1e-5 barely trains in 10 epochs (loss 3.393 with icing, 3.349 without), while lr 3e-3 reaches 2.199 with a constant schedule and 2.063 with warmup-cosine. No held-out PPL was measured.

  • lr 1e-5, entropy icing on, 10 epochs: train loss (nats, training trie) = 3.393104
  • lr 3e-3 warmup-cosine, 10 epochs: train loss (nats, training trie) = 2.063377

Sources. Epoch loss lines and header settings in train.log, lr3e-3/train.log, lr3e-3-wc/train.log and no-icing/train.log; commit 4a5e479 message calls these LR probe runs.

Caveats. rnd/TRIAGE.md lists it under REMOVE as a one-off LR probe. v1 CUDA trainer, before the 2026-05-26 loss fix. The headline values come from train.log, not from a README or result.json.