← All experiments THE RESEARCH RECORD / heldout-tree-vs-model

Held-out trie vs trained model

Does the trie alone (count lookup with backoff) predict held-out text as well as a trained AGPT model, i.e. is the trie doing the real work?

experimentconcludednegativeattentioneval: legacy
Opened Updated
THE ANSWER SO FAR

No. On a 10% contiguous held-out tail of Gutenberg 5M, the 100-SE d64 L2 model scores legacy PPL 5.03 against 170.39 for naive-backoff trie lookup. Naive backoff overstates the 34x ratio, but the ordering is robust. A pd=1 shuffle ablation gave 5.29 (single seed, not distinguishable).

trained model, 100 SE, pd=1 5.03 legacy held-out PPL (4096 positions, seq 16, Gutenberg 10% tail) pre-loss-fixpre-race-fixtruncated-ancestor-gradientlegacy-eval
trie alone, naive backoff 170.39 legacy held-out PPL (4096 positions, seq 16, Gutenberg 10% tail) pre-loss-fixpre-race-fixtruncated-ancestor-gradientlegacy-eval
trained model, pd=1 with shuffle 5.29 legacy held-out PPL (4096 positions, seq 16, Gutenberg 10% tail) pre-loss-fixpre-race-fixtruncated-ancestor-gradientlegacy-eval
TOPICS
RELATED
SUPERSEDED BY
CODE
THE FULL RECORD

Experiment notes

Open directory on GitHub ↗

Held-out trie vs trained model

Stub README (2026-09-28): this directory predates the README convention. The summary below was reconstructed from its files; see them for detail.

Question. Does the trie alone (count lookup with backoff) predict held-out text as well as a trained AGPT model, i.e. is the trie doing the real work?

Answer. No. On a 10% contiguous held-out tail of Gutenberg 5M, the 100-SE d64 L2 model scores legacy PPL 5.03 against 170.39 for naive-backoff trie lookup. Naive backoff overstates the 34x ratio, but the ordering is robust. A pd=1 shuffle ablation gave 5.29 (single seed, not distinguishable).

  • trained model, 100 SE, pd=1: legacy held-out PPL (4096 positions, seq 16, Gutenberg 10% tail) = 5.03
  • trie alone, naive backoff: legacy held-out PPL (4096 positions, seq 16, Gutenberg 10% tail) = 170.39
  • trained model, pd=1 with shuffle: legacy held-out PPL (4096 positions, seq 16, Gutenberg 10% tail) = 5.29

Sources. findings.md (Question, Result table, Interpretation, Caveats, Shuffle ablation); logs/model_ppl_heldout.log (5.0274) and logs/trie_ppl_heldout.log (170.3934).

Caveats. Outcome 'negative' refers to the tested hypothesis 'the trie does the work'. The finding that the trie does not generalize holds only for the naive-backoff evaluator, as the doc's own caveat 4 says. The later count-backoff-gate learned gate reaches heldout fixed PPL ~3.9 on Shakespeare. Model and trie artifacts lived in /tmp.