Vanilla attention LM (PyTorch)
How does the AGPT attention architecture (d64 L2 h4 ff256) perform when trained as a plain mini-batch Adam language model, without the trie, on the canonical held-out evaluation?
Explore the tests, diagnostics, baselines, designs, and infrastructure behind AGPT. Each page preserves its protocol and links to the full record.
98 published records · attention, recurrent, count-prior, and hybrid work
Filter the research record.
How does the AGPT attention architecture (d64 L2 h4 ff256) perform when trained as a plain mini-batch Adam language model, without the trie, on the canonical held-out evaluation?
Does giving attention extra K/V slots for suffix-backoff trie nodes (B=4, gradient through the anc-grad path) lower held-out PPL at Shakespeare d=16?
Can cheaper recurrent f_theta (linear, linear+RMSNorm, GRU and GRU variants) replace attention inside AGPT's trie training, and how close do they come?
What does the population of per-unit gradients at a frozen model reveal about the cost of prefix sharing (fewer optimizer updates per epoch), and can a better cadence or optimizer recover that cost?
With the CUDA v2 attention trainer, do alternatives to all-depth pd=1 training (whole-trie pd=0 updates, single-depth loss, or a deterministic pd=6 descendant sweep) match or beat the pd=1 baseline?
Can giving each prefix node a calibrated family of skip-horizon empirical distributions (targets h>1 steps ahead, each with its own backoff ladder) feed AGPT long-range signal without breaking the aggregated-gradient identity?
Can a small recurrent f_θ (tanh recurrence) inside AGPT, later extended with a neural history residual on top of a count/backoff prior, give a strong character model on Tiny Shakespeare?
Does the count-prior + segment-memory residual transfer to a 5M-character Gutenberg corpus, and what does the residual need in order to add information beyond a strong prior?
Does a neural residual over segment memory (carried GRU or gated cross-attention), trained on top of a frozen recursive count-gate prior, improve held-out PPL beyond the prior alone?
Can trie-derived variable-length segments serve as a route vocabulary for a recurrent character model with attention over previous segment states, and how should local recurrence and segment memory be combined?
Does a product-of-experts backoff prior (log p_root plus gated log p_d along drop-oldest suffix chains, with a 5-parameter logistic gate) make a usable trie prior?
Can sampling many clustered mini-trees from the full trie (one optimizer step each) recover stochastic update cadence and beat the pd1 whole-tree AGPT baseline?
Can a pd1 whole-tree AGPT checkpoint be improved by a short stochastic mini-tree refinement phase using the same attention model?
In a matched segment harness, does a zero-initialized gated cross-attention block over segment memory records improve on a carried GRU, and does the answer depend on data scale?
Can a count-only model learn from local prefix statistics when to trust a deeper context instead of backing off, and do its smoothed distributions help as neural AGPT targets?
Do prefix trees built over strided corpus positions (stride 2, 4, 16) expose longer-range structure that complements the adjacent-character tree?
Can trie-aggregated Fisher natural-gradient updates (an exact merged output-head Fisher step, optionally a body Fisher step) train a small GRU character model competitively, over the whole trie or over bounded prefix subtrees?
Can a trie-derived prior (backoff counts, or a state-Fisher-corrected head) plus a learned neural residual improve held-out prediction when the prior is built from training counts only?
Do node-local natural steps on trie hidden states, using a state-space Fisher built from each node's count row, fit the trie objective, and can the corrected states be carried into a shared GRU (or a node-state table) that generalizes?
Do child-to-parent messages in the trie (Fisher- or precision-weighted merges of child evidence, pseudo-counts, logit models or adapters) improve on purely local per-node fitting?
Can AGPT's prefix-trie aggregation train a non-attention f_theta, specifically a simple tanh recurrence over prefix transitions?
Can a depth-16 AGPT path see further by appending the real observed successor cap path as extra attention context?
Can presenting trie paths at RoPE positions/phases that match where their prefixes occur in the corpus improve the depth-16 CUDAX AGPT recipe?
How do we run AGPT experiments on rented RunPod GPUs without hand-maintaining the environment?
Can a prebuilt container image give RunPod pods a ready AGPT build environment without per-pod setup?
Which partition depth, optimizer, learning rate and epoch budget give the best static full-trie CUDAX v2 AGPT baseline at depth 16 on Shakespeare, scored on the sampled multi-chunk heldout?
Does a GRU encoder over the 16 characters before a node's prefix, residual-injected at layer 0, extend AGPT's effective context and lower held-out PPL?
What held-out PPL does a standard sliding-window microgpt model (d64 L2, seq_len 16, 10k steps) reach under the AGPT experiment-harness split and evaluator?
How does CUDAX progressive-growth AGPT at several division × epoch schedules compare with a µGPT sliding-window SGD baseline of the same model size and context length?
Does feeding a trie node's predecessors' deepest hidden state (h_cap, mass-weighted over corpus predecessors) into the node as a side input h_in extend AGPT's context and lower loss at Shakespeare d=16?
How does the static v2 d64/L2 calibration setup score on a contiguous 5% tail held-out split, compared with the sampled multi-chunk split?
What held-out PPL does a standard sliding-window microgpt model reach at seq_len 124 (d64 L2, about 25 corpus epochs), as a reference for depth-124 AGPT?
How does the v1 trainer compare with v2 under the canonical evaluator, and how does v1 respond to mass weighting (linear/sqrt/log)?
Does adding the asymmetric-DFT harmonic-filter bias to attention scores improve held-out perplexity of a small char-level LM on Shakespeare?
If the rolling-PPL irregularity in CUDAX progressive-growth runs comes from progressive staging, do static full-tree runs improve monotonically with epochs?
Under the YAML harness with a tail held-out split, does giving the 16-division progressive-growth schedule more epochs per frontier improve held-out byte PPL?
What tail-heldout PPL does the v2 trainer's d=16 trie, d64/L2 static-epoch baseline reach with paper-correct linear mass weighting at 25, 100 and 250 epochs?
Can the CUDAX v2 trainer train a d64 L2 model on the full depth-124 Shakespeare trie, and how does held-out PPL move with the number of static epochs?
What tail-heldout PPL does a standard window-trained transformer (d128 L6, seq_len=16, 180k steps) reach at roughly the wall time of CUDAX d128/L6 static200?
Can the CUDAX v2 trainer handle a full depth-124 Shakespeare radix trie (the depth at which raw contexts become unique), and if not, what is the blocker?
What char-level KenLM KN PPL does Shakespeare reach by n-gram order on a 10-chunk random 95/5 disjoint held-out split?
With the Section 2 event-count weighting restored in the CUDAX (v2) trainer, what held-out PPL do progressive-growth schedules (divisions × epochs) and static full-prefix training reach on Shakespeare?
What held-out byte perplexity does the standard small AGPT recipe (d64 L2, depth-16 trie, 10 static epochs, rmsprop 3e-3, anc-grad) reach on Shakespeare's 5% tail?
Does CUDAX static full-prefix training behave correctly after restoring radix exposure and switching to Section 2 (paper) event-count weighting?
Do position-distribution 'chord' features of trie substrings separate on-path from off-path query/key pairs, and which operator and frequency set separates them best?
At a fixed budget of 100 super-epochs on Gutenberg 5M, does widening d_model or adding layers push AGPT held-out PPL past the L=6 d=64 ceiling?
Does replacing chunk-local RoPE positions with per-substring position-distribution summaries (dist-rope, or the scalar expected position) help AGPT training?
Does training the CUDAX trainer on progressively growing corpus prefixes (N growth divisions times epochs per stage, depth 16) improve held-out PPL, and which growth schedule works best?
Where does char-level Kneser-Ney (KenLM) PPL plateau by n-gram order on Gutenberg 5M, and can per-trie-node KN distributions be extracted as soft targets?
On a provably disjoint Gutenberg held-out set, can AGPT at moderate scale beat a classical Kneser-Ney character model?
With anc-grad on, does any per-event weighting flag (mass, entropy, branching or depth weight, five modes each) improve Gutenberg 5M PPL, and does Shakespeare's branching=log win transfer?
Does AGPT use RoPE as a literal sequence coordinate, or can trie depth be swapped for another monotonic signal (edge mass, log mass) without hurting PPL?
Does multiplying mass and entropy per-event loss weights together carry over to Gutenberg, where the single-axis Shakespeare wins did not?
Is RMSprop's slow beta2 transient the bottleneck in short AGPT runs, so that beta2=0.99 at 10 SE matches beta2=0.999 at 100 SE?
What held-out PPL does the legacy v1 trainer reach on Shakespeare 1M (d=16, 10 SE, anc-grad off) from the shared seed models, as a parity reference for v2?
Does weighting per-event loss by a single axis (mass, depth, entropy or branching), under three fire-end normalization regimes, improve held-out PPL?
Does the v2 (CUDAX) trainer's descendant-to-ancestor Wk/Wv gradient (anc-grad) reproduce the legacy trainer's held-out improvement over anc-grad off?
How does the new v2 (CUDAX) trainer train on a standard pd=1 run, for comparison with the v1 trainer?
Does training while the trie grows (successive corpus-prefix checkpoints, carrying model and optimizer state) beat one-shot full-trie AGPT at a matched SE budget?
Should weight gradients be normalized once per optimizer fire (1/N events) instead of per memory chunk (1/T_q_chunk)?
How does 10-epoch RMSProp training on the Shakespeare d=16 radix trie compare at lr 1e-5 (with and without branching-endpoint entropy icing) against lr 3e-3 (constant and warmup-cosine)?
Does restoring descendant-to-ancestor gradient flow into Wk/Wv (--anc-grad, with the per-event normalizer fix) improve held-out PPL?
Does localizing the optimizer second-moment state to root-child subtree buckets (per-rc Adam/RMSprop) improve AGPT training over one global state?
Does streaming AGPT (100 stages x 5 SE) beat single-stage 500-SE training on Gutenberg 5M at depth 16, as it did on Shakespeare?
Does the trie alone (count lookup with backoff) predict held-out text as well as a trained AGPT model, i.e. is the trie doing the real work?
Can pooling predictions or activations from overlapping d=16 windows at inference time beat the d=16 model's own perplexity on Gutenberg 5M without retraining?
Can a d=16-trained AGPT model use context beyond its trie depth through naive RoPE extrapolation, and can the per-position trie-node bookkeeping needed for shared-key RoPE be built?
Do next-token distributions read directly from the forward radix trie equal those recovered by Bayesian inversion of the suffix (reversed-corpus) trie?
Is microgpt's cuBLAS training path correct, and how much wall-clock speedup does it give over openBLAS?
How does partition depth (optimizer fires per super-epoch) trade training quality against wall time on Gutenberg 5M at trie depth 16/18?
Does replacing the one-hot targets at the first cap-tunnel positions with composite shifted-prefix distributions improve AGPT PPL at d=32?
Does replacing the unary-tunnel walk with a structural wormhole jump (cap to depth-1 re-entry node) give better synthetic-corpus training signal than synth_wrap's walk-and-bridge?
Does a stop-gradient KL consistency loss between a forward (prefix) model and a backward (reversed-suffix) model shrink their divergence and improve forward-only PPL?
Does replacing one-hot radix-cap targets with the corpus-wide suffix distribution P(c|W) improve PPL on Shakespeare d=32?
Do partition depth, progressive curriculum and hotspot coverage compose when stacked, since each adds optimizer-fire density?
Does firing one optimizer step per depth-N prefix group (--partition-depth N --no-accumulate) instead of one per root child speed up AGPT convergence?
Does splitting the radix trie into a root-side decision zone and a leaf-side identity zone (K = decision, V = identity) predict trie statistics and AGPT depth behaviour, and can it be turned into a training improvement?
Does randomly dropping root-child subtrees each super-epoch improve AGPT training through trajectory variety or dropout-like regularization?
Was AGPT undertrained at the standard 3 super-epoch budget, i.e. does PPL keep falling with more super-epochs using the same recipe?
Can structural max-overlap matching between prefix-trie and suffix-trie leaves serve as an attention or prediction backbone that beats a direct transformer on next-character PPL?
Can a mass cap plus ancestor K/V warmup make Lightning L3 sampling on Gutenberg 5M (depth 32) match L4 perplexity at much lower wall-clock cost?
Does the wrap-around synthetic-corpus pipeline scale from Shakespeare to a 5M-character Gutenberg corpus?
Does a depth-D radix trie carry enough predictive content to train models with seq_len > D? Tested by training on a corpus synthesized from trie walks that wrap from leaf back to root.
Do mass-1 unary chains in the d=32 radix trie carry useful training signal, or can they be pruned from the synthetic wrap-around corpus without cost?
Do AGPT and standard SGD over corpus positions reach similar held-out PPL at matched compute, and which AGPT mass weighting (off, log, sqrt, linear) matches SGD?
How much of AGPT's perplexity advantage over plain SGD on Shakespeare comes from subtree aggregation, and how much from optimizer and recipe differences?
Does splitting the root-child subtrees that carry the most excess loss between epochs improve AGPT PPL over a uniform per-root-child sweep?
Does AGPT subtree training need an adaptive optimizer, or can plain SGD or momentum match RMSProp?
After commit 1c858c0 made Wk, Wv and the biases trainable, which learning rate and schedule give the best deterministic AGPT baseline at d=16 and d=32, and how does it compare with sliding-window SGD?
Can stochastic variable-depth subtree sampling (Lightning L3 mass-weighted walk) match or beat the deterministic per-root-child sweep at the same number of optimizer steps?
Does converged PPL track the radix-trie saturation curve, giving diminishing gains from d=8 to 16 to 32 and none past 32?
Can the CUDA trainer size its KV cache to the largest partition group rather than the whole trie file, so that --partition-depth reduces peak memory enough to train the full d=16 global trie in one super-epoch?
How do node count, branching and singleton fraction change with depth in the Shakespeare radix tries (d=8, 16, 32)?
Does training over a virtual tree of depth K·D, built by stitching copies of the D-trie at its leaves, improve PPL over K=1?
Does the trie depth cap distort path-probability convergence beyond the ordinary noise of mass-1 paths?
Do path-probability products from tries built on two independent halves of a corpus agree, and how does agreement change with trie depth?
Does blending shorter-suffix count distributions into the target at radix endpoints (count-aware smoothing) improve AGPT PPL?
Does packing deep unary-chain segments into one variable-length forward per group (a two-regime trainer: sibling-grouped shallow depths, packed deep chains) speed up the Crystal trie-walk trainer without changing the update cadence?
No experiments match these filters.