State-Fisher geometry
Freeze the GRU and its head, treat each trie node's hidden state as a free variable, and take
damped natural steps delta_h_v = -(F_v + damping I)^-1 g_v with the state-space Fisher
F_v = N_v W^T (Diag(r_v) - r_v r_v^T) W built from the node's count row, plus an exact
per-node line search in logit space. Then try to move the improvement into a shared model.
Results (Tiny Shakespeare, stride-1 circular samples, block 8, d64, damping 10, empirical curvature; legacy aggregate train-trie PPL, no held-out): with no parameter update the trie PPL fell from 63.52 to 7.70 (1 iteration), 4.187 (8) and 3.994 (16). Projection into the GRU was weak (legacy sequential validation PPL, 90/10 contiguous split): corrected-logit distillation 23.12 with a frozen head (10 epochs) and 9.18 with a trainable head (5 epochs); state-delta projection 28.72; hidden-target distillation with a merged Fisher head 13.25 after five passes (whole-tree variant 14.04). The pure book state model (a free node-state table plus one Fisher-updated head, longest-suffix lookup at eval) reached train PPL 4.94 by epoch 20, while held-out PPL bottomed at 25.14 (epoch 6) and rose to 31.50.
Conclusion: the count-derived local Fisher geometry is real, but one shared nonlinear model does not absorb the free node-state moves, and free node states overfit.
Code and records (under research/ultra/):
agpt_ultra/state_fisher.py,state_projection.py,fisher.pyscripts/run_state_fisher_diagnostic.py,run_state_distill.py(--distill-target hidden),run_state_delta_project.py,run_book_state_model.py,generate_book_state.pynotebook/archive/state_fisher_results.md: "Setup" through "Projection Notes", hidden-target rows of "Return To AGPT Training", "Pure Book State Model"docs/trie_fisher_bridge.md,docs/trie_fisher_bridge_revised.md(theory)