Tree-prior residual
Model the logits as a trie-derived prior plus a learned residual, z_final = z_prior + z_residual. Three priors were tried: a suffix/backoff count prior, a detached
state-Fisher-corrected head prior, and the explicit log-count backoff distribution with forced
suffix drops during training. Held-out evaluation uses train-derived priors only (longest
train-seen suffix), never held-out counts.
Results (Tiny Shakespeare, contiguous last 10% held out; legacy PPL): the backoff prior with a d128 residual moved held-out 10.123 to 9.995 (depth 8). The Fisher-prior residual looked strong on the train trie (block 16: 2.166 -> 2.076, at 37-40 min per epoch and about 10 GB RSS) but scored 12.782 held-out (Fisher prior alone 13.243). On the first 2k validation samples the backed-off count prior (6.601) beat the Fisher prior (9.395) and Fisher + residual (9.137), and head-only fidelity training barely moved it (9.366). With the explicit count prior, a conservative depth-8 residual (scale 0.25, LR 3e-4) improved held-out 6.191 to 6.157.
Conclusion: the state/head projection loses the count evidence, so the Fisher prior does not
generalize; an explicit count prior plus residual behaves as expected but gains little. The
strand was dropped as not competitive with ordinary SGD on a small attention model; count
priors returned later as the recursive count gate (see count-prior-residual).
Code and records (under research/ultra/):
scripts/run_residual_prior.py,run_fisher_residual.py,eval_fisher_residual.py,run_prior_fidelity.py,run_direct_prior_residual.pyagpt_ultra/state_fisher.py,flat_trie.pynotebook/archive/state_fisher_results.md: "Residual Tree Prior", "Fisher-Prior Residual", "Held-Out Evaluation", "Direct Tree-Prior Residual"