Independent research · 2026

One epoch
and done.

What if a language model could learn enough from one pass over a corpus to rival conventional training?

AGPT (Aggregated Gradient Pre-training) explores a different way to organize training: share work across repeated text prefixes, aggregate their evidence, and ask more of each update.

One epoch is an ideal, probably out of reach. This research asks how close we can get.

THE QUESTIONCan a better training structure make each corpus pass count more?
01 / THE IDEA

Stop recomputing
the same beginning.

Text windows often share prefixes. A trie gives those shared beginnings one path, then branches when the text differs.

THREE TOKENS, TWO VIEWS

The same next-token examples, trained two ways.

SGD and AGPT views of three next-token examples AGPT stores the cas prefix once in a trie before branching to the next tokens t, e, and h. cas●teh ONE SHARED PATHTHREE BRANCHES

AGPT represents the shared cas path once. Each branch contributes evidence back through that shared computation.

THE MATHEMATICAL MOVE

For a shared prefix, sum descendant gradient contributions before applying the common prefix transformation.

J · Σgs = Σ(J · gs)

This is an exact algebraic identity. Turning it into a correct, fast transformer training system is the hard part.

Read the paper draft
02 / THE ADVANTAGE

Reuse the work.
Keep the pattern.

When text repeats, a tree can turn that repetition into shared computation and a reusable map of the corpus.

01 / COMPUTE

Pay for a prefix once.

Separate examples may calculate the same prefix state many times. AGPT can share that state across branches and, ideally, compute it once per corpus pass. A node may cost more to process, but fewer repeated evaluations could reduce total work.

Potential gainFewer repeated calculationsTradeoffTracking cached states
02 / STRUCTURE

Repetition shapes the tree.

Repeated contexts share one path. In Tiny Shakespeare, branching nodes peak at 78,836 at character depth 7, then fall to 6,119 by depth 16. At that cap, 99.43% of depth-16 contexts have one observed continuation: zero empirical next-token entropy. A radix trie folds one-child stretches into edges; other tree structures might compress them further.

16.2×node-count compression versus an uncompressed trie at depth 32 on the 1.1M-character Shakespeare corpus
THE BRANCHING STRUCTURE PLATEAUS

On the same 1.1M-character corpus, going from depth 16 to 124 adds just 3.7% more radix node records:

1.608Mdepth 161.666Mdepth 321.668Mdepth 124

Open scale question: Could branching across all recorded human writing end within roughly 50 tokens? This is a working estimate, still unmeasured. Storage at that scale would depend on how the tails are represented.

By depth 124, Tiny Shakespeare has no repeated character contexts left. Radix edges already compress the long one-child paths; pruning tails or looping back through the root may reduce storage further. Those ideas and any end-to-end speed advantage still need testing. Branching profile ↗ · Depth-124 profile ↗ · Tail pruning ↗ · Root wrap ↗

03 / THE CHALLENGES

Three problems
to solve.

Shared-prefix training has to find useful context beyond the tree, update often enough to learn, and pay for its extra computation.

01CONTEXT DEPTH

Useful context outruns the reusable tree.

Sharing works because short contexts repeat. Long ones rarely do: extend a context far enough and nearly every path through the tree has only one observed continuation. At depth 16 in Tiny Shakespeare, over 99% do. Yet language depends on context far longer than the point where repetition runs out.

WHAT SUCCESS LOOKS LIKE

Long-range context and prefix sharing in the same model, rather than one traded for the other.

See how we’ve approached this
02UPDATE CADENCE

More sharing means fewer updates.

Every branch below a shared prefix feeds one combined step. That is where the savings come from, and also the cost: the weights move once where ordinary training would move them many times. When branches pull in different directions, one combined step can learn less than many smaller ones would have.

WHAT SUCCESS LOOKS LIKE

Keeping the savings of sharing while learning at least as much per unit of computation as ordinary training.

See how we’ve approached this
03COMPUTE COST

Sharing has to pay for itself.

Ordinary training spends its compute on many small, noisy updates. A shared pass spends it on fewer updates, each built from exact information about the whole corpus. That trade wins only if the exact information is worth more than the updates given up, and only if maintaining the tree costs less than the work it saves: the caches, the gradient through shared prefixes, the structure itself.

WHAT SUCCESS LOOKS LIKE

Matched model quality in less wall-clock time than conventional training, on the same hardware, counting every cost.

See how we’ve approached this
04 / THE FRAMEWORK

It began with GPT.
The idea grew.

AGPT started as a way to rethink GPT training. Later, the same structure turned out to apply to other models with reusable prefix states.

WHAT MAKES A MODEL FIT?

If the state after a prefix contains everything needed to process the next token, the same prefix can be computed once and reused across its branches. With differentiable transitions and additive losses, descendant gradients can be combined before passing through the shared computation.

hp·x = fθ(hp, x)

The tracks used different evaluation procedures. Their historical perplexity numbers are documented separately and are not directly comparable.

05 / THE NEXT TESTS

Move the ball
forward.

“One epoch and done” is the measuring stick. The next steps test whether better gradients and better updates can close the gap.

  1. 01

    Complete the CUDA backward pass

    Make the aggregated gradient match finite differences through ancestor attention.

  2. 02

    Run matched evaluations

    Keep corpus splits, context, evaluator, model size, and wall time explicit.

  3. 03

    Test curvature on the exact trie gradient

    Find out whether one expensive, information-rich pass can replace many noisy updates.

  4. 04

    Measure the one-pass gap

    Report how close one corpus pass gets in quality and time, even when it falls short.

THE RESEARCH IS OPEN

Curious about the details?

The code, paper draft, run records, and dead ends are public. Start with the research record and follow the evidence.