Partition-scoped KV cache
The CUDA trainer sizes its KV cache for a whole trie file, even when
--partition-depth splits the file into groups that train one after another. At d=16
the global trie needs about 9.5 GB of KV and does not fit. This branch planned to
size the cache to the largest partition group instead. Each group would get a CPU-side
global_to_local index remap and an ancestor mini-forward, with no kernel changes.
Only Phase 1 was built: a --partition-kv-scoped flag that prints the achievable
savings and does not change behaviour. For the largest d=16 per-subtree file with
bigram groups it reported peak KV of 1295.7 MB unscoped vs 161.7 MB scoped (8.0x).
Phase 2 (the scoped allocation, remap, mini-forward, a parity test and the d=16
global run) is fully specified in the plan but was never implemented.
Code: branch agpt-partition-kv-scoping, tag exp/partition-kv-scoping. Key files
are notes/agpt/partition-kv-scoping-plan.md (the full Phase 2 plan) and
src/cuda/agpt_train.cu (Phase 1 stats, commit 48ee729).