← All experiments THE RESEARCH RECORD / vanilla-attn-torch

Vanilla attention LM (PyTorch)

How does the AGPT attention architecture (d64 L2 h4 ff256) perform when trained as a plain mini-batch Adam language model, without the trie, on the canonical held-out evaluation?

baselineconcludedn/aattentioneval: canonical
Opened Updated
THE ANSWER SO FAR

It beats AGPT at this scale. From random init at seq 16 it reaches canonical byte PPL 4.46 by epoch 10 and plateaus at 4.45, versus 4.85 for the canonical AGPT run (depth-16 trie, seed, 512 epochs, wrap). The trainer is src/tools/agpt_vanilla_attn_train.py on tag exp/linear-recurrence-final; the run outputs were not kept, so the numbers come from the 2026-06-08 project notes.

TOPICS
RELATED
SUPERSEDED BY
CODE
THE FULL RECORD

Experiment notes

Open directory on GitHub ↗

Vanilla attention LM (PyTorch)

Stub README (2026-09-28): this directory predates the README convention. The summary below was reconstructed from its files; see them for detail.

Question. How does the AGPT attention architecture (d64 L2 h4 ff256) perform when trained as a plain mini-batch Adam language model, without the trie, on the canonical held-out evaluation?

Answer. It beats AGPT at this scale. From random init at seq 16 it reaches canonical byte PPL 4.46 by epoch 10 and plateaus at 4.45, versus 4.85 for the canonical AGPT run (depth-16 trie, seed, 512 epochs, wrap). This is recorded only in the 2026-06-08 memory note; the run directories and trainer are not in the repository. (The run files were not kept; the numbers come from the 2026-06-08 project notes.)

Sources. Memory note project_vanilla_attn_beats_agpt.md (table of vanilla SGD vs canonical AGPT). The directory itself holds no results.

Caveats. The directory is gitignored/untracked and contains only d64L2-seq32-scratch/train.log, a failed launch ('can't open file src/tools/agpt_vanilla_attn_train.py'). That trainer has since been committed on branch worktree-linear-recurrence (tag exp/linear-recurrence-final, src/tools/agpt_vanilla_attn_train.py, 2026-09-28). Memory also lists seq=32 ep50 at 4.17 on the same held-out, but the only surviving seq32 artefact is this failed launch, so that number's provenance needs checking. The headline values come from memory, not from a README or result.