Product-of-experts backoff prior
A Python prototype of a suffix-link backoff prior: exact tuple-key contexts from the training
text, drop-oldest backoff chains (ABCD -> BCD -> CD -> D -> root), and a product of experts in
logit space, logits = log p_root + sum_d gate(features_d) * log p_d, with fixed count tensors
and only a 5-parameter logistic gate trainable. It tested the math before any radix or suffix
catalog was involved.
Results (Tiny Shakespeare, 200k-char setting, prior only, legacy validation PPL): root 28.812, deepest context 13.840, uniform product 224.823, conservatively initialized gate (bias -2) 10.130. Entropy damping helped deepest-only (13.014 at a=0.5) and the uniform product (48.991 at a=1.0) but worsened the initialized gate (11.949 at a=0.5). Training the gate overfit even at LR 0.005: validation went 10.130 -> 10.917 -> 11.394 -> 12.361 while train PPL fell toward 1.
Conclusion: multiplying sharp count distributions is unsafe, and a learned gate memorizes train
statistics. The count-gate reproduction that followed (count-backoff-gate) uses a recursive
convex mixture, q_d = w_d p_mle_d + (1 - w_d) q_backoff, and the archive says the product form
likely explains this prototype's instability.
Code and records (under research/ultra/):
scripts/run_backoff_gate_prior.pynotebook/archive/state_fisher_results.md: "Backoff-Gated Trie Prior", and the contrast drawn in "Count Backoff Gate Reproduction"