THE RESEARCH RECORD

Every experiment.
One place.

Explore the tests, diagnostics, baselines, designs, and infrastructure behind AGPT. Each page preserves its protocol and links to the full record.

98 published records · attention, recurrent, count-prior, and hybrid work

FIND AN EXPERIMENT

Filter the research record.

98 experimentsClick a title for the full record
rnd/vanilla-attn-torch

Vanilla attention LM (PyTorch)

How does the AGPT attention architecture (d64 L2 h4 ff256) perform when trained as a plain mini-batch Adam language model, without the trie, on the canonical held-out evaluation?

n/aattention
Opened 2026-06-08
Updated 2026-09-28
rnd/slot-selection-step0

Backoff slot selection

Does giving attention extra K/V slots for suffix-backoff trie nodes (B=4, gradient through the anc-grad path) lower held-out PPL at Shakespeare d=16?

negativeattention
Opened 2026-05-30
Updated 2026-09-28
rnd/linear-recurrence

Linear and GRU recurrence

Can cheaper recurrent f_theta (linear, linear+RMSNorm, GRU and GRU variants) replace attention inside AGPT's trie training, and how close do they come?

mixedrecurrent
Opened 2026-06-06
Updated 2026-09-28
rnd/gradient-population

Gradient population

What does the population of per-unit gradients at a frozen model reveal about the cost of prefix sharing (fewer optimizer updates per epoch), and can a better cadence or optimizer recover that cost?

n/aattention
Opened 2026-09-24
Updated 2026-09-28
rnd/stochastic-agpt

Stochastic AGPT (v2 cadence and depth controls)

With the CUDA v2 attention trainer, do alternatives to all-depth pd=1 training (whole-trie pd=0 updates, single-depth loss, or a deterministic pd=6 descendant sweep) match or beat the pd=1 baseline?

negativeattention
Opened 2026-06-12
Updated 2026-09-27
rnd/skip-horizon-agpt

Skip-horizon AGPT

Can giving each prefix node a calibrated family of skip-horizon empirical distributions (targets h>1 steps ahead, each with its own backoff ladder) feed AGPT long-range signal without breaking the aggregated-gradient identity?

n/an/a
Opened 2026-06-18
Updated 2026-06-18
rnd/rnn-agpt

RNN AGPT

Can a small recurrent f_θ (tanh recurrence) inside AGPT, later extended with a neural history residual on top of a count/backoff prior, give a strong character model on Tiny Shakespeare?

negativerecurrent
Opened 2026-06-12
Updated 2026-06-14
rnd/gutenberg-prior-residual

Gutenberg prior residual

Does the count-prior + segment-memory residual transfer to a 5M-character Gutenberg corpus, and what does the residual need in order to add information beyond a strong prior?

inconclusivehybrid
Opened 2026-06-13
Updated 2026-06-14
rnd/count-prior-residual

Count-prior residual

Does a neural residual over segment memory (carried GRU or gated cross-attention), trained on top of a frozen recursive count-gate prior, improve held-out PPL beyond the prior alone?

mixedhybrid
Opened 2026-06-12
Updated 2026-06-13
rnd/segment-memory

Segment memory

Can trie-derived variable-length segments serve as a route vocabulary for a recurrent character model with attention over previous segment states, and how should local recurrence and segment memory be combined?

mixedrecurrent
Opened 2026-06-11
Updated 2026-06-12
rnd/poe-backoff-prior

Product-of-experts backoff prior

Does a product-of-experts backoff prior (log p_root plus gated log p_d along drop-oldest suffix chains, with a 5-parameter logistic gate) make a usable trie prior?

negativecount-prior
Opened 2026-06-12
Updated 2026-06-12
rnd/lightning-agpt

Lightning AGPT

Can sampling many clustered mini-trees from the full trie (one optimizer step each) recover stochastic update cadence and beat the pd1 whole-tree AGPT baseline?

mixedattention
Opened 2026-06-12
Updated 2026-06-12
rnd/hybrid-agpt

Hybrid pd1 + stochastic AGPT

Can a pd1 whole-tree AGPT checkpoint be improved by a short stochastic mini-tree refinement phase using the same attention model?

negativeattention
Opened 2026-06-12
Updated 2026-06-12
rnd/gated-xattn-memory

Gated cross-attention memory

In a matched segment harness, does a zero-initialized gated cross-attention block over segment memory records improve on a carried GRU, and does the answer depend on data scale?

mixedrecurrent
Opened 2026-06-12
Updated 2026-06-12
rnd/count-backoff-gate

Count backoff gate

Can a count-only model learn from local prefix statistics when to trust a deeper context instead of backing off, and do its smoothed distributions help as neural AGPT targets?

mixedcount-prior
Opened 2026-06-12
Updated 2026-06-12
rnd/stride-trees

Stride trees

Do prefix trees built over strided corpus positions (stride 2, 4, 16) expose longer-range structure that complements the adjacent-character tree?

inconclusiverecurrent
Opened 2026-06-11
Updated 2026-06-11
rnd/trie-fisher-gru

Trie-Fisher GRU training

Can trie-aggregated Fisher natural-gradient updates (an exact merged output-head Fisher step, optionally a body Fisher step) train a small GRU character model competitively, over the whole trie or over bounded prefix subtrees?

negativerecurrent
Opened 2026-06-09
Updated 2026-06-10
rnd/tree-prior-residual

Tree-prior residual

Can a trie-derived prior (backoff counts, or a state-Fisher-corrected head) plus a learned neural residual improve held-out prediction when the prior is built from training counts only?

mixedhybrid
Opened 2026-06-10
Updated 2026-06-10
rnd/state-fisher-geometry

State-Fisher geometry

Do node-local natural steps on trie hidden states, using a state-space Fisher built from each node's count row, fit the trie objective, and can the corrected states be carried into a shared GRU (or a node-state table) that generalizes?

mixedrecurrent
Opened 2026-06-09
Updated 2026-06-10
rnd/model-reconciliation

Model reconciliation

Do child-to-parent messages in the trie (Fisher- or precision-weighted merges of child evidence, pseudo-counts, logit models or adapters) improve on purely local per-node fitting?

negativerecurrent
Opened 2026-06-10
Updated 2026-06-10
rnd/tanh-recurrence

Tanh recurrent AGPT

Can AGPT's prefix-trie aggregation train a non-attention f_theta, specifically a simple tanh recurrence over prefix transitions?

positiverecurrent
Opened 2026-06-07
Updated 2026-06-07
rnd/successor-prefix-attention

Successor prefix attention

Can a depth-16 AGPT path see further by appending the real observed successor cap path as extra attention context?

inconclusiveattention
Opened 2026-06-05
Updated 2026-06-06
rnd/sampled-node-phase-rope

Sampled-node phase RoPE

Can presenting trie paths at RoPE positions/phases that match where their prefixes occur in the corpus improve the depth-16 CUDAX AGPT recipe?

positiveattention
Opened 2026-06-05
Updated 2026-06-05
rnd/runpod

RunPod launcher

How do we run AGPT experiments on rented RunPod GPUs without hand-maintaining the environment?

n/an/a
Opened 2026-05-17
Updated 2026-06-05
rnd/docker

AGPT Docker image

Can a prebuilt container image give RunPod pods a ready AGPT build environment without per-pod setup?

n/an/a
Opened 2026-05-18
Updated 2026-06-05
rnd/baseline-calibration-v2-static-sampled

v2 static calibration (sampled heldout)

Which partition depth, optimizer, learning rate and epoch budget give the best static full-trie CUDAX v2 AGPT baseline at depth 16 on Shakespeare, scored on the sampled multi-chunk heldout?

n/aattention
Opened 2026-05-30
Updated 2026-06-05
rnd/precondition

Precondition encoder

Does a GRU encoder over the 16 characters before a node's prefix, residual-injected at layer 0, extend AGPT's effective context and lower held-out PPL?

negativeattention
Opened 2026-06-01
Updated 2026-06-02
rnd/shake-sgd-baseline

Shakespeare SGD window baseline

What held-out PPL does a standard sliding-window microgpt model (d64 L2, seq_len 16, 10k steps) reach under the AGPT experiment-harness split and evaluator?

n/aattention
Opened 2026-05-30
Updated 2026-05-30
rnd/progressive-growth-sgd-comparison

Progressive growth vs SGD

How does CUDAX progressive-growth AGPT at several division × epoch schedules compare with a µGPT sliding-window SGD baseline of the same model size and context length?

inconclusiveattention
Opened 2026-05-26
Updated 2026-05-30
rnd/cap-recurrence

Cap recurrence

Does feeding a trie node's predecessors' deepest hidden state (h_cap, mass-weighted over corpus predecessors) into the node as a side input h_in extend AGPT's context and lower loss at Shakespeare d=16?

negativeattention
Opened 2026-05-27
Updated 2026-05-30
rnd/baseline-calibration-v2-static-tail

v2 static baseline, tail split

How does the static v2 d64/L2 calibration setup score on a contiguous 5% tail held-out split, compared with the sampled multi-chunk split?

n/aattention
Opened 2026-05-30
Updated 2026-05-30
rnd/window-d124-baseline

Window baseline at seq_len 124

What held-out PPL does a standard sliding-window microgpt model reach at seq_len 124 (d64 L2, about 25 corpus epochs), as a reference for depth-124 AGPT?

n/aattention
Opened 2026-05-28
Updated 2026-05-29
rnd/v1-vs-v2-comparison

v1 vs v2 trainer comparison

How does the v1 trainer compare with v2 under the canonical evaluator, and how does v1 respond to mass weighting (linear/sqrt/log)?

inconclusiveattention
Opened 2026-05-26
Updated 2026-05-29
rnd/harmonic-bias-prototype

Harmonic-filter attention bias prototype

Does adding the asymmetric-DFT harmonic-filter bias to attention scores improve held-out perplexity of a small char-level LM on Shakespeare?

negativeattention
Opened 2026-05-26
Updated 2026-05-29
rnd/cudax-static-epochs

CUDAX static epoch sweep

If the rolling-PPL irregularity in CUDAX progressive-growth runs comes from progressive staging, do static full-tree runs improve monotonically with epochs?

positiveattention
Opened 2026-05-26
Updated 2026-05-29
rnd/cudax-growth-heldout-rerun

CUDAX progressive-growth held-out rerun

Under the YAML harness with a tail held-out split, does giving the 16-division progressive-growth schedule more epochs per frontier improve held-out byte PPL?

mixedattention
Opened 2026-05-29
Updated 2026-05-29
rnd/cudax-d16-linear-mass-rerun

CUDAX d16 linear-mass rerun

What tail-heldout PPL does the v2 trainer's d=16 trie, d64/L2 static-epoch baseline reach with paper-correct linear mass weighting at 25, 100 and 250 epochs?

n/aattention
Opened 2026-05-28
Updated 2026-05-29
rnd/cudax-d124-probe

CUDAX depth-124 probe

Can the CUDAX v2 trainer train a d64 L2 model on the full depth-124 Shakespeare trie, and how does held-out PPL move with the number of static epochs?

inconclusiveattention
Opened 2026-05-28
Updated 2026-05-29
rnd/window-baseline-wallmatch

Window baseline, wall-matched

What tail-heldout PPL does a standard window-trained transformer (d128 L6, seq_len=16, 180k steps) reach at roughly the wall time of CUDAX d128/L6 static200?

n/aattention
Opened 2026-05-28
Updated 2026-05-28
rnd/radix-depth124

Depth-124 radix trie

Can the CUDAX v2 trainer handle a full depth-124 Shakespeare radix trie (the depth at which raw contexts become unique), and if not, what is the blocker?

n/an/a
Opened 2026-05-28
Updated 2026-05-28
rnd/kn-shakespeare-baseline

KN Shakespeare multi-chunk baseline

What char-level KenLM KN PPL does Shakespeare reach by n-gram order on a 10-chunk random 95/5 disjoint held-out split?

n/acount-prior
Opened 2026-05-28
Updated 2026-05-28
rnd/cudax-section2-progressive

CUDAX Section 2 progressive reruns

With the Section 2 event-count weighting restored in the CUDAX (v2) trainer, what held-out PPL do progressive-growth schedules (divisions × epochs) and static full-prefix training reach on Shakespeare?

n/aattention
Opened 2026-05-26
Updated 2026-05-27
rnd/shake-small-baseline

Small Shakespeare baseline

What held-out byte perplexity does the standard small AGPT recipe (d64 L2, depth-16 trie, 10 static epochs, rmsprop 3e-3, anc-grad) reach on Shakespeare's 5% tail?

n/aattention
Opened 2026-05-25
Updated 2026-05-26
rnd/cudax-exposure-fix

CUDAX exposure fix

Does CUDAX static full-prefix training behave correctly after restoring radix exposure and switching to Section 2 (paper) event-count weighting?

n/aattention
Opened 2026-05-26
Updated 2026-05-26
rnd/harmonic-filter-diagnostic

Harmonic filter diagnostic

Do position-distribution 'chord' features of trie substrings separate on-path from off-path query/key pairs, and which operator and frequency set separates them best?

n/an/a
Opened 2026-05-25
Updated 2026-05-25
rnd/dmodel-scaling-100se

d_model and depth scaling at 100 SE (Gutenberg)

At a fixed budget of 100 super-epochs on Gutenberg 5M, does widening d_model or adding layers push AGPT held-out PPL past the L=6 d=64 ceiling?

mixedattention
Opened 2026-05-24
Updated 2026-05-25
rnd/dist-rope-smoke

dist-rope smoke test

Does replacing chunk-local RoPE positions with per-substring position-distribution summaries (dist-rope, or the scalar expected position) help AGPT training?

negativeattention
Opened 2026-05-25
Updated 2026-05-25
rnd/cudax-growth

CUDAX progressive growth

Does training the CUDAX trainer on progressively growing corpus prefixes (N growth divisions times epochs per stage, depth 16) improve held-out PPL, and which growth schedule works best?

mixedattention
Opened 2026-05-25
Updated 2026-05-25
rnd/kenlm-baseline

KenLM KN baseline

Where does char-level Kneser-Ney (KenLM) PPL plateau by n-gram order on Gutenberg 5M, and can per-trie-node KN distributions be extracted as soft targets?

n/acount-prior
Opened 2026-05-24
Updated 2026-05-24
rnd/scale-vs-kn

AGPT scaling vs Kneser-Ney

On a provably disjoint Gutenberg held-out set, can AGPT at moderate scale beat a classical Kneser-Ney character model?

positiveattention
Opened 2026-05-23
Updated 2026-05-23
rnd/gutenberg-anc-sweep

Gutenberg weighting sweep with anc-grad

With anc-grad on, does any per-event weighting flag (mass, entropy, branching or depth weight, five modes each) improve Gutenberg 5M PPL, and does Shakespeare's branching=log win transfer?

negativeattention
Opened 2026-05-23
Updated 2026-05-23
rnd/rope-position-substitution

RoPE position substitution

Does AGPT use RoPE as a literal sequence coordinate, or can trie depth be swapped for another monotonic signal (edge mass, log mass) without hurting PPL?

positiveattention
Opened 2026-05-22
Updated 2026-05-22
rnd/composite-weights

Composite mass × entropy weights

Does multiplying mass and entropy per-event loss weights together carry over to Gutenberg, where the single-axis Shakespeare wins did not?

negativeattention
Opened 2026-05-21
Updated 2026-05-22
rnd/beta2-diagnostic

RMSprop beta2 vs training length

Is RMSprop's slow beta2 transient the bottleneck in short AGPT runs, so that beta2=0.99 at 10 SE matches beta2=0.999 at 100 SE?

negativeattention
Opened 2026-05-21
Updated 2026-05-22
rnd/legacy-rebaseline

Legacy v1 rebaseline

What held-out PPL does the legacy v1 trainer reach on Shakespeare 1M (d=16, 10 SE, anc-grad off) from the shared seed models, as a parity reference for v2?

n/aattention
Opened 2026-05-21
Updated 2026-05-21
rnd/depth-weight

Per-event loss weighting

Does weighting per-event loss by a single axis (mass, depth, entropy or branching), under three fire-end normalization regimes, improve held-out PPL?

negativeattention
Opened 2026-05-21
Updated 2026-05-21
rnd/cudax-anc-grad-parity

CUDAX anc-grad parity

Does the v2 (CUDAX) trainer's descendant-to-ancestor Wk/Wv gradient (anc-grad) reproduce the legacy trainer's held-out improvement over anc-grad off?

n/aattention
Opened 2026-05-21
Updated 2026-05-21
rnd/v2-compare

v2 trainer comparison run

How does the new v2 (CUDAX) trainer train on a standard pd=1 run, for comparison with the v1 trainer?

n/aattention
Opened 2026-05-20
Updated 2026-05-20
rnd/streaming-agpt-v1

Streaming AGPT

Does training while the trie grows (successive corpus-prefix checkpoints, carrying model and optimizer state) beat one-shot full-trie AGPT at a matched SE budget?

positiveattention
Opened 2026-05-16
Updated 2026-05-20
rnd/per-fire-norm

Per-fire gradient normalization

Should weight gradients be normalized once per optimizer fire (1/N events) instead of per memory chunk (1/T_q_chunk)?

mixedattention
Opened 2026-05-20
Updated 2026-05-20
rnd/lr1e-5-probe

LR 1e-5 probe

How does 10-epoch RMSProp training on the Shakespeare d=16 radix trie compare at lr 1e-5 (with and without branching-endpoint entropy icing) against lr 3e-3 (constant and warmup-cosine)?

n/aattention
Opened 2026-05-20
Updated 2026-05-20
rnd/anc-grad

Ancestor gradient (anc-grad)

Does restoring descendant-to-ancestor gradient flow into Wk/Wv (--anc-grad, with the per-event normalizer fix) improve held-out PPL?

positiveattention
Opened 2026-05-20
Updated 2026-05-20
rnd/per-rc-adam-v1

Per-root-child Adam state

Does localizing the optimizer second-moment state to root-child subtree buckets (per-rc Adam/RMSprop) improve AGPT training over one global state?

negativeattention
Opened 2026-05-18
Updated 2026-05-18
rnd/overnight-2026-05-18

Streaming AGPT on Gutenberg (overnight run)

Does streaming AGPT (100 stages x 5 SE) beat single-stage 500-SE training on Gutenberg 5M at depth 16, as it did on Shakespeare?

positiveattention
Opened 2026-05-18
Updated 2026-05-18
rnd/heldout-tree-vs-model

Held-out trie vs trained model

Does the trie alone (count lookup with backoff) predict held-out text as well as a trained AGPT model, i.e. is the trie doing the real work?

negativeattention
Opened 2026-05-18
Updated 2026-05-18
rnd/sliding-window-v1

Sliding-window AGPT v1

Can pooling predictions or activations from overlapping d=16 windows at inference time beat the d=16 model's own perplexity on Gutenberg 5M without retraining?

negativeattention
Opened 2026-05-11
Updated 2026-05-11
rnd/seq-len-decouple

Seq_len decoupling

Can a d=16-trained AGPT model use context beyond its trie depth through naive RoPE extrapolation, and can the per-position trie-node bookkeeping needed for shared-key RoPE be built?

negativeattention
Opened 2026-05-11
Updated 2026-05-11
rnd/prefix-suffix-bayes

Prefix-suffix Bayesian consistency

Do next-token distributions read directly from the forward radix trie equal those recovered by Bayesian inversion of the suffix (reversed-corpus) trie?

n/an/a
Opened 2026-05-11
Updated 2026-05-11
rnd/microgpt-cublas-verify

microgpt cuBLAS verification

Is microgpt's cuBLAS training path correct, and how much wall-clock speedup does it give over openBLAS?

n/aattention
Opened 2026-05-10
Updated 2026-05-11
rnd/gutenberg-pd-sweep

Gutenberg 5M partition-depth sweep

How does partition depth (optimizer fires per super-epoch) trade training quality against wall time on Gutenberg 5M at trie depth 16/18?

mixedattention
Opened 2026-05-11
Updated 2026-05-11
rnd/virtual-tree

Virtual tree (composite cap targets)

Does replacing the one-hot targets at the first cap-tunnel positions with composite shifted-prefix distributions improve AGPT PPL at d=32?

negativeattention
Opened 2026-05-06
Updated 2026-05-07
rnd/wormhole

Wormhole routing

Does replacing the unary-tunnel walk with a structural wormhole jump (cap to depth-1 re-entry node) give better synthetic-corpus training signal than synth_wrap's walk-and-bridge?

negativeattention
Opened 2026-05-06
Updated 2026-05-06
rnd/dual-model-fold

Dual-view consistency (forward + backward models)

Does a stop-gradient KL consistency loss between a forward (prefix) model and a backward (reversed-suffix) model shrink their divergence and improve forward-only PPL?

inconclusiveattention
Opened 2026-05-05
Updated 2026-05-05
rnd/cap-folding

Cap folding

Does replacing one-hot radix-cap targets with the corpus-wide suffix distribution P(c|W) improve PPL on Shakespeare d=32?

mixedattention
Opened 2026-05-04
Updated 2026-05-05
rnd/granularity-redundancy

Granularity redundancy

Do partition depth, progressive curriculum and hotspot coverage compose when stacked, since each adds optimizer-fire density?

negativeattention
Opened 2026-05-01
Updated 2026-05-01
rnd/partition-depth

Partition depth

Does firing one optimizer step per depth-N prefix group (--partition-depth N --no-accumulate) instead of one per root child speed up AGPT convergence?

positiveattention
Opened 2026-04-30
Updated 2026-04-30
rnd/trie-attention-framing

Trie-as-attention framing

Does splitting the radix trie into a root-side decision zone and a leaf-side identity zone (K = decision, V = identity) predict trie statistics and AGPT depth behaviour, and can it be turned into a training improvement?

mixedattention
Opened 2026-04-28
Updated 2026-04-29
rnd/subtree-dropout

Subtree dropout

Does randomly dropping root-child subtrees each super-epoch improve AGPT training through trajectory variety or dropout-like regularization?

negativeattention
Opened 2026-04-29
Updated 2026-04-29
rnd/agpt-epoch-scaling

AGPT epoch scaling

Was AGPT undertrained at the standard 3 super-epoch budget, i.e. does PPL keep falling with more super-epochs using the same recipe?

positiveattention
Opened 2026-04-29
Updated 2026-04-29
rnd/p2s-attention

Prefix-to-suffix attention

Can structural max-overlap matching between prefix-trie and suffix-trie leaves serve as an attention or prediction backbone that beats a direct transformer on next-character PPL?

negativeattention
Opened 2026-04-27
Updated 2026-04-27
rnd/lightning-cap-warmup

Lightning L3 mass cap and ancestor warmup

Can a mass cap plus ancestor K/V warmup make Lightning L3 sampling on Gutenberg 5M (depth 32) match L4 perplexity at much lower wall-clock cost?

mixedattention
Opened 2026-04-27
Updated 2026-04-27
rnd/gutenberg-5m

Gutenberg 5M wrap-around

Does the wrap-around synthetic-corpus pipeline scale from Shakespeare to a 5M-character Gutenberg corpus?

positiveattention
Opened 2026-04-26
Updated 2026-04-26
rnd/wrap-around

Wrap-around corpus synthesis

Does a depth-D radix trie carry enough predictive content to train models with seq_len > D? Tested by training on a corpus synthesized from trie walks that wrap from leaf back to root.

positiveattention
Opened 2026-04-25
Updated 2026-04-25
rnd/unary-pruning

Mass-1 unary-path pruning

Do mass-1 unary chains in the d=32 radix trie carry useful training signal, or can they be pruned from the synthetic wrap-around corpus without cost?

mixedattention
Opened 2026-04-25
Updated 2026-04-25
rnd/sgd-sanity-check

SGD vs AGPT sanity check

Do AGPT and standard SGD over corpus positions reach similar held-out PPL at matched compute, and which AGPT mass weighting (off, log, sqrt, linear) matches SGD?

inconclusiveattention
Opened 2026-04-21
Updated 2026-04-25
rnd/sgd-ceiling

SGD ceiling

How much of AGPT's perplexity advantage over plain SGD on Shakespeare comes from subtree aggregation, and how much from optimizer and recipe differences?

mixedattention
Opened 2026-04-24
Updated 2026-04-25
rnd/hotspot-curriculum

Hotspot curriculum

Does splitting the root-child subtrees that carry the most excess loss between epochs improve AGPT PPL over a uniform per-root-child sweep?

mixedattention
Opened 2026-04-24
Updated 2026-04-25
rnd/agpt-optimizers

AGPT optimizers

Does AGPT subtree training need an adaptive optimizer, or can plain SGD or momentum match RMSProp?

positiveattention
Opened 2026-04-25
Updated 2026-04-25
rnd/post-fix-baseline

Post Wk/Wv-fix baseline

After commit 1c858c0 made Wk, Wv and the biases trainable, which learning rate and schedule give the best deterministic AGPT baseline at d=16 and d=32, and how does it compare with sliding-window SGD?

n/aattention
Opened 2026-04-23
Updated 2026-04-24
rnd/lightning-training

Lightning training

Can stochastic variable-depth subtree sampling (Lightning L3 mass-weighted walk) match or beat the deterministic per-root-child sweep at the same number of optimizer steps?

negativeattention
Opened 2026-04-22
Updated 2026-04-23
rnd/radix-saturation

Radix saturation vs PPL

Does converged PPL track the radix-trie saturation curve, giving diminishing gains from d=8 to 16 to 32 and none past 32?

inconclusiveattention
Opened 2026-04-22
Updated 2026-04-22
rnd/partition-kv-scoping

Partition-scoped KV cache

Can the CUDA trainer size its KV cache to the largest partition group rather than the whole trie file, so that --partition-depth reduces peak memory enough to train the full d=16 global trie in one super-epoch?

n/an/a
Opened 2026-04-22
Updated 2026-04-22
rnd/sparsity-profile

Trie sparsity profile

How do node count, branching and singleton fraction change with depth in the Shakespeare radix tries (d=8, 16, 32)?

n/an/a
Opened 2026-04-21
Updated 2026-04-21
rnd/root-loop

Root-loop virtual tree

Does training over a virtual tree of depth K·D, built by stitching copies of the D-trie at its leaves, improve PPL over K=1?

negativeattention
Opened 2026-04-21
Updated 2026-04-21
rnd/mass-conservation

Mass conservation and depth cap

Does the trie depth cap distort path-probability convergence beyond the ordinary noise of mass-1 paths?

n/an/a
Opened 2026-04-21
Updated 2026-04-21
rnd/convergence

Trie path-probability convergence

Do path-probability products from tries built on two independent halves of a corpus agree, and how does agreement change with trie depth?

n/an/a
Opened 2026-04-18
Updated 2026-04-21
rnd/blending

Suffix-depth blending

Does blending shorter-suffix count distributions into the target at radix endpoints (count-aware smoothing) improve AGPT PPL?

mixedattention
Opened 2026-04-21
Updated 2026-04-21
rnd/packed-varlen

Packed variable-length forward

Does packing deep unary-chain segments into one variable-length forward per group (a two-regime trainer: sibling-grouped shallow depths, packed deep chains) speed up the Crystal trie-walk trainer without changing the update cadence?

negativeattention
Opened 2026-04-15
Updated 2026-04-15