Gated cross-attention memory
Follow-up to segment-memory. First a controlled check of the segment harness: at 200k train
chars, d64, LR 1e-3 and about 3,860 updates, a random-window sequence GRU reached 7.83 and the
carried segment-stream GRU 7.757 after 10 epochs (first 20k validation chars), so the harness
GRU floor is not handicapped. Then a Flamingo-style block (--mixing gated-xattn),
h + tanh(g_attn) * MHA(LN(h), LN(mem), LN(mem)) plus a gated MLP, with both gates
zero-initialized so training starts exactly at the carried GRU.
Results (legacy segment-harness held-out PPL, first 20k chars of the tail-10% split): on the 200k slice the block fell behind the no-memory GRU by epoch 2 (epoch 3: 9.575 vs 8.577); a terminal-record auxiliary at weight 0.1 helped slightly (9.458) and 0.25 did not. On the full 1,003,854-char training split (108,041 segments) it led at every epoch: 6.503 vs 6.988 after one epoch and 5.199 vs 5.627 after five, still descending. Cost: about 2.2x the no-memory epoch time on the full corpus and about 2.5x on 200k (backward is about 62% of train time; checkpointing is negligible).
Conclusion: route-memory attention adds modeling capacity once there are enough routes; the
50k/200k slices were underpowered. This configuration became the residual model on top of the
frozen count prior (count-prior-residual).
Code and records (under research/ultra/):
scripts/run_segment_memory_model.py(--mixing gated-xattn,--memory-record written,--terminal-record-aux-weight,--max-memory,--profile,--checkpoint-output,--resume-checkpoint)notebook/archive/state_fisher_results.md: "Matched GRU Harness Check And Gated Cross-Attention", "Full-Corpus Segment-Memory Check", "Segment-Memory Checkpointing And Profile"notebook/archive/segment_memory_math.md: "Controlled Harness Check" through "Full-Corpus Route Scale"