← All experiments THE RESEARCH RECORD / trie-fisher-gru

Trie-Fisher GRU training

Can trie-aggregated Fisher natural-gradient updates (an exact merged output-head Fisher step, optionally a body Fisher step) train a small GRU character model competitively, over the whole trie or over bounded prefix subtrees?

experimentconcludednegativerecurrenteval: legacy
Opened Updated
THE ANSWER SO FAR

No, not in these prototypes. On bounded depth-8 subtrees an AdamW GRU body with a merged Fisher head reached legacy held-out sequential PPL 11.96 after five passes (15.76 after one); body Fisher on all prefixes diverged to 2287.49 after one pass, and per-prefix head Fisher on partial evidence drove held-out PPL to infinity. Sequence training of the same GRU family reached validation PPL 5.78 (context 8, 50,000 updates), and a small AdamW attention LM reached held-out window PPL 5.600 (seq len 16, 5,000 steps).

AdamW GRU body + merged Fisher head, depth-8 subtrees, 5 passes 11.96 legacy held-out sequential PPL (Tiny Shakespeare, script-default 90/10 contiguous split) pre-race-fixlegacy-eval
Body Fisher GRU + Fisher head, all prefixes, 1 pass 2287.49 legacy held-out sequential PPL (same protocol) pre-race-fixlegacy-eval
Reference: sequence-trained GRU, context 8, 50,000 updates 5.78 legacy validation PPL pre-race-fixlegacy-eval
TOPICS
RELATED
SUPERSEDED BY
CODE
THE FULL RECORD

Experiment notes

Open directory on GitHub ↗

Trie-Fisher GRU training

The first AGPT Ultra prototypes trained a small GRU character model (Embedding -> GRUCell -> linear head) on the trie objective, updating the output head with an exact, matrix-free trie Fisher natural step, theta -= (sum_p F_p + lambda I)^-1 sum_p g_p solved by conjugate gradient, and the body with ordinary gradient steps (AdamW in the subtree runs) or a body-Fisher step. Later runs used bounded depth-8 prefix subtrees as the training units.

Results (Tiny Shakespeare, script-default contiguous 90/10 split, legacy held-out sequential PPL): AdamW body + merged Fisher head reached 15.76 after one pass and 11.96 after five. Body Fisher on the a/e/t prefixes reached 16.16 at about 9 GB RSS; on all prefixes it diverged to 2287.49 after one pass. Per-prefix head Fisher on partial evidence drove held-out PPL to infinity, so head evidence has to be merged after a full sweep or guarded. For reference, sequence training of the same GRU family reached validation PPL 6.27 (5,000 updates) and 5.78 (50,000 updates) at context 8, and a 2-layer AdamW attention LM reached held-out window PPL 5.600 (seq len 16, 5,000 steps). An earlier prefix-subtree run is recorded only as "roughly 8.22 to 8.26 PPL at depth 8". Conclusion: the merged head Fisher step is stable, but training the GRU this way was not competitive with ordinary sequence training.

Code and records (under research/ultra/):