Motivation
My dream when I started learning about LLMs through Andrej Karpathy's video series was to train the best language model available. That dream was quickly abandoned after I discovered how expensive even training a competitive model with 1B parameters would be!
That led me to look for ways to improve training efficiency by reading and implementing research papers. After almost 300 runs and experiments, I quite fortunately came up with a modification that became ExoFormer: Attention Projection Mixing with Exogenous Anchors, which was accepted into the main track of ICML 2026.
But I never truly gave up on the goal of training a state of the art language model. After discovering the Open SLM Leaderboard by Axiomic Labs, I was delighted to find categories for extremely small models. I wanted to try my hand at a model under 3M parameters first, then explore larger sizes later.
Why so small?
I've always wanted to train an absurdly small language model on an extremely large number of tokens. At 100B tokens and roughly 3M parameters, that is about 33,000 tokens per parameter: over 1,600 times the familiar Chinchilla rule of thumb of about 20.
Would the training loss just flatten out? Would there be any sign of grokking, or would improvements remain gradual?
I've also wanted to try attention residuals in recurrent language models. Reusing the same block keeps the parameter count small, while a learned residual mixture gives each loop more freedom to select information from earlier loops. It is also refreshing to be able to iterate and test ablations so quickly with a model this small.
Much of the training code came from modded nanogpt. I left out its U Net residuals, value residual learning (sadly), and large n gram embedding tables. Here is the final architecture:
- A vocabulary of 2,048 tokens with tied input and output embeddings.
- One shared block applied six times: hidden width 384 and FFN width 2,096.
- Fused ReLU² in the backbone.
- LR AttnRes with rank 128.
- CE + NITP, using a SwiGLU projector and auxiliary weight 1.
- NorMuon for the backbone matrices and Adam for embeddings and small parameters.
- Six attention heads, a context of 8,192 tokens, and BF16 training.
- Per head XSA and token dependent sigmoid gates on the attention outputs.
Each head gets an XSA coefficient and a sigmoid gate from a 16 channel controller. XSA adjusts the component of the attention output aligned with the current value vector, then the gate scales the head before the output projection. Together, these add 192 parameters.
Vocabulary is expensive at this scale
My intuition is that a large vocabulary is a particularly unproductive use of parameters in a model this small. With width 384, even a tied input/output vocabulary of 2,048 tokens costs 786,432 parameters, roughly a quarter of the model. Increasing that to 8,192 tokens would cost 3,145,728 parameters for the embedding table alone. That already exceeds the entire budget.
That is why I kept the vocabulary of 2,048 tokens. There is a tradeoff: smaller vocabularies can require more tokens to represent the same text, spending more sequence length and computation.
At width 384, a tied vocabulary of 2,048 tokens uses 786,432 parameters. A vocabulary of 8,192 uses 3,145,728 parameters, exceeding the full 3M budget.
Residuals across recurrent loops
I started with LR AttnRes, my low rank attention residual method, using my Fast AttnRes implementation. It learns how to mix information from earlier loops for each token. In the model with three loops, the routing queries add just 768 parameters.
I compared it with ordinary pre norm residual addition and a learned scalar baseline. The scalar baseline learns fixed weights for earlier residual sources. LR AttnRes chooses its mixture based on the token's representations.
Three ways to read the residual stream
Training CE · nats/token · Final displayed points| LR AttnRes | 2.5997 |
|---|
| Learned scalar | 2.6799 |
|---|
| Pre norm | 2.6956 |
|---|
LR AttnRes gave me lower loss in both views. At three loops, final heldout loss was 2.6239, compared with 2.7152 for learned scalar residuals and 2.7336 for pre norm. Training took about 11% longer than the scalar baseline and 17% longer than pre norm.
Why I moved from three loops to six
My initial loop sweep used about 1B tokens. The extra loops helped, but the gains became small beyond six: heldout loss was 2.6239 at three loops, 2.6116 at six, 2.6156 at nine, and 2.6101 at twelve. Twelve was marginally best at equal tokens, while three processed tokens about 3.9 times faster.
How much recurrence is enough?
Training CE · nats/token · Final displayed points| 1 loop | 2.8660 |
|---|
| 2 loops | 2.6331 |
|---|
| 3 loops | 2.5844 |
|---|
| 6 loops | 2.5623 |
|---|
| 9 loops | 2.5637 |
|---|
| 12 loops | 2.5603 |
|---|
I also swept the loop count for learned scalar residuals and pre norm. Adding loops helped differently depending on how I connected them.
Residual routing across depth
Heldout CE · nats/token · Final displayed points| LR AttnRes | 2.6101 |
|---|
| Learned scalar | 2.6605 |
|---|
| Pre norm | 2.6706 |
|---|
I later tested three and six loops over a longer training horizon. Six loops reached lower training loss within the same recorded training time window, despite processing fewer tokens per second. At 12B tokens, its heldout loss was 2.2260, compared with 2.2993 for three loops.
Six loops over a longer horizon
Training CE · nats/token · Final displayed points| Three loops | 2.3261 |
|---|
| Six loops | 2.2755 |
|---|
I also finished a longer run with twelve loops. At 12B tokens its heldout loss was lower again, 2.1979, but its length normalized local intelligence index was 5.758, below the 6.039 reached by the six loop model trained with CE alone.
NITP was worth revisiting
Next Implicit Token Prediction, or NITP, adds a second objective alongside cross entropy for next token prediction. I use a small projection head to predict the next token's shallow representation from the current token's final hidden state. In my version, that target comes from the first recurrent block.
I like the idea of learning from a target in hidden space as well as a token ID! The paper stopped gradients through the target, so the auxiliary loss cannot improve simply by moving the target toward its prediction.
At three loops, the early NITP screen did not improve heldout loss. But the result changed at six loops. With the SwiGLU projector and auxiliary weight 1, the 12B heldout loss moved from 2.2993 to 2.3037 at three loops, and from 2.2260 to 2.2172 at six loops. The direction was the same at 10B.
NITP changes with depth
Heldout CE · nats/token · Displayed values| 3 loops · CE only | 2.3260 |
|---|
| 3 loops · CE + NITP | 2.3301 |
|---|
| 6 loops · CE only | 2.2548 |
|---|
| 6 loops · CE + NITP | 2.2448 |
|---|
My intuition is that more depth gives the hidden states more room to transform between the shallow target and the final prediction. I haven't isolated that explanation. The projector also mattered: the linear projector at six loops reached 2.2320, worse than the CE only control. I kept CE + NITP with the SwiGLU projector.
My six loop model has 2,987,712 inference parameters. The NITP head adds 1,769,472 during training, bringing the total to 4,757,184. I discard the head for inference, so the model I use afterward is still under 3M parameters.
What happens to NITP during cooldown?
I also tested whether to change NITP during cooldown. From the same checkpoint with six loops and a linear projector at 10B tokens, I trained four separate continuations to 12B: freeze the projector, remove NITP, keep the projector's learning rate constant, or keep NITP and reduce the learning rate floor to zero. I restarted each branch from the same full checkpoint.
Changing the objective late in training
Heldout CE · nats/token · Displayed values| Original · LR floor 0.1 · Original | 2.2320 |
|---|
| Frozen projector · Frozen head | 2.2322 |
|---|
| CE only cooldown · CE only | 2.2244 |
|---|
| Constant projector LR · Fixed head LR | 2.2321 |
|---|
| NITP · LR floor 0 · LR floor 0 | 2.2305 |
|---|
Removing NITP during cooldown gave the lowest heldout CE, 2.2244, versus 2.2320 for the original cooldown. It did not give the best downstream index: 6.029 versus 6.128. Freezing the projector barely changed heldout CE (2.2322) while raising that index to 6.346. The branch with a constant projector learning rate was also nearly unchanged on CE.
I found the disagreement between loss and downstream scores more interesting than the small ranking differences. I evaluated these older runs in BF16, and changed deterministic execution settings between the original run and the continuations. With one seed and that difference, I wasn't convinced freezing the projector was better. I kept the later SwiGLU projector trainable.
DeepCrossAttention did not make the cut
I also tried DeepCrossAttention, another way of reading earlier depth representations. In my short adaptation, heldout loss rose from 2.5885 to 2.6230, and counted training took 32.6% longer than its matched LR AttnRes control.
DeepCrossAttention against its control
Training CE · nats/token · Final displayed points| LR AttnRes | 2.5459 |
|---|
| DeepCrossAttention | 2.6345 |
|---|
I did not keep it after this test over roughly 1B tokens. I only tried one seed and my own unfused adaptation.
For learned residual routing, I focused on AttnRes and DeepCrossAttention. I haven't tried HC, iHC, or mHC in this model yet. I'd like to explore those later.
One large shared block or two smaller blocks?
I tried spending the budget on two distinct blocks instead of one larger shared block. To fit them, I reduced the hidden width from 384 to 256. I compared the same number of block applications: two blocks repeated three times execute six blocks, just like one block repeated six times.
One wide block, or two narrower blocks?
Training CE · nats/token · Final displayed points| 1 block × 6 loops · width 384 | 2.2902 |
|---|
| 2 blocks × 3 loops · width 256 | 2.3170 |
|---|
At six block applications, the single wider block scaled better over the shared token window. Around 4.24B tokens, its smoothed training CE was 2.2902, versus 2.3170 for two smaller blocks. That made “one big block” an attractive direction.
At twelve block applications, the result was more mixed. Around 2.61B tokens, the two block model had slightly lower smoothed loss: 2.2996 versus 2.3122. I couldn't call one wider block universally better from that. Twelve applications already felt too slow for my budget, though, and I wanted to try the two pass idea below, so I stayed with one wider block.
Small details that mattered
BF16 worked better than the FP8 paths I tried. Moving the FFN and output projection to BF16 reduced heldout loss from 2.5889 to 2.5013, while training time fell from 297.2 to 292.2 seconds. Changing the output projection made the biggest difference to loss. I suspect conversion overhead made FP8 less useful at this size.
Precision at a very small scale
Training CE · nats/token · Final displayed points| FP8 FFN + output | 2.5379 |
|---|
| BF16 FFN · FP8 output | 2.5283 |
|---|
| BF16 throughout | 2.4534 |
|---|
SwiGLU did worse than fused ReLU² in my backbone. At the same parameter count, I got 2.5137 heldout loss with SwiGLU versus 2.5085 with ReLU², and training took 6% longer. I had to change the FFN width and its optimizer settings to fit the budget, so I couldn't attribute the difference to the activation alone. I still used SwiGLU in the separate NITP head.
ReLU² and SwiGLU in the backbone
Training CE · nats/token · Final displayed points| ReLU² · width 2,112 | 2.4613 |
|---|
| SwiGLU · width 1,408 | 2.4685 |
|---|
Using every last parameter was not automatically better. In three paired BF16 seeds, FFN width 2,096 beat 2,112 on heldout loss each time: mean loss 2.5003 versus 2.5064. I kept 2,096 although it was really confusing.
The harsh lesson from 100B tokens
I learned how much care the data deserves. Before the long run, I investigated repeated rises and falls in training loss. Replaying the loader traced 38 of 39 selected excursions to documents drawn from the FineMath 4+ stream in my assembled corpus. Some contained long base64 image payloads, notebook fragments, extraction debris, or repetitive text. Popular dataset names had made me too comfortable with what was actually reaching the model.
My first completed run with 100B tokens was a harsh lesson in a different way. I had spent plenty of effort on architecture while choosing the data mixture too casually. The length normalized local intelligence index stayed below 6 at every 10B milestone, ending at 5.681. Meanwhile, heldout loss continued falling, from 2.1347 at 10B to 2.0613 at 100B.
A hundred billion tokens later
Local intelligence index · points · Final displayed points| Length normalized | 5.6810 |
|---|
What really frustrated me was watching heldout loss improve without sustained gains in downstream accuracy.
Trying the Pulvis v2 architecture
I wanted to compare my wide shared block with a narrower model containing more distinct blocks. I adapted the layout of Pulvis v2, while keeping my tokenizer, AttnRes, NITP, optimizer, and batch size. I trained both on the same token stream: 4B tokens at a constant learning rate, then 1B of linear cooldown to zero.
My six loop model used one block of width 384 applied six times. The model based on Pulvis v2 used ten distinct blocks of width 160 across sixteen applications, repeating its middle three blocks three times. Attention, FFN, and context length also differed.
I changed quite a lot at once, so I was comparing the two architectures as a whole. I couldn't isolate the effect of weight sharing or claim to have reproduced the published Pulvis model.
Six loop uses one shared block of width 384 applied six times. The Pulvis v2 shape uses ten stored blocks of width 160 across sixteen applications.
Pulvis v2 shape versus six loops
Heldout CE · nats/token · Displayed values| Six loop · Six loop | 2.0232 |
|---|
| Pulvis v2 shaped · Pulvis v2 shape | 2.1071 |
|---|
After 5B tokens, the model based on Pulvis v2 had the higher length normalized local intelligence index: 6.673 versus 6.172. The six loop model had substantially better heldout CE: 2.0232 versus 2.1071. The task results were mixed: the narrower architecture gained 2.02 percentage points on ARC Easy and 1.09 on PIQA, while losing 1.45 on ARC Challenge and 0.48 on HellaSwag. ArithMark3 was almost unchanged.
The narrower adaptation also used fewer parameters: 2,637,376 at inference and 2,944,576 during training, compared with 2,987,712 and 4,757,184 for six loops. Both models in my comparison used a vocabulary of 2,048 tokens. The published Pulvis v2 configuration uses 4,096. NITP's head scales with width, so that difference in training parameters matters.
I evaluated both completed runs on the full heldout split and the same downstream suite in FP32. I only used one seed, but I found the split between lower loss and higher downstream scores interesting.
Latent tokens and a simpler two pass model
PonderLM 2 gave me another direction to explore: reuse a token's final hidden state as a latent input before predicting the next token. Its sequential latent dependencies make naive parallel training difficult. Jacobi iterations approximate those dependencies by refining all positions in parallel.
My simpler variant keeps the initial and final passes, with zero Jacobi iterations:
Six loops build a hidden state for each real token. Interleave each state with its real token, then apply six loops again. The second pass latent outputs predict the next real token.
The second pass has twice as many positions because I interleave the real tokens with latent inputs. Across both passes, I process L + 2L positions for L real training tokens.
Using the paper's approximate work formula, 3 + 2K times a single pass for K Jacobi rounds, I get 3× at zero Jacobi, compared with 7× or 9× for two or three rounds. Those are estimates of training work, not measured GPU times.
In an earlier three loop pilot, I sampled two or three Jacobi rounds per update. Three round updates took a median 0.955 seconds, versus 0.334 seconds with zero Jacobi. The plot compares the run mixing two and three Jacobi rounds with zero Jacobi over the same first 1B tokens.
Jacobi refinement versus zero Jacobi
Training CE · trailing 100 updates · Final displayed points| Jacobi · 2 to 3 rounds | 2.3441 |
|---|
| Zero Jacobi | 2.3517 |
|---|
That is a lot of repeated work for my budget, which is why the plain two pass version appealed to me.
For the six loop version, each pass executes six loops through the shared backbone, with an independent bank of learned depth queries for each pass. CE and NITP supervise the latent outputs from the second pass. The inference parameter count is 2,989,248, and the training total including the NITP head is 4,758,720.
The timing view below uses H100 GPU hours, calculated from recorded training cycle time and GPU count. Both curves cover the same hardware budget during the plateau phase, before cooldown begins. The timer includes checkpointing and evaluation within training cycles, but excludes startup and unrecorded restart gaps.
Two passes: tokens and compute tell different stories
Training CE · trailing 200 updates · Final displayed points| Single pass | 2.0210 |
|---|
| Two passes | 2.0901 |
|---|
I ultimately abandoned two pass training because it was too time consuming. Even without Jacobi iterations, the extra computation took more time than I could justify for this project, so I returned to the single pass six loop model.
The final model
My best saved model is the single pass six loop version, with 2,987,712 inference parameters and a local intelligence index of 8.4873. I trained it for 4B tokens at a plateau learning rate followed by 1B of cooldown, on a heavily filtered mixture similar to Pulvis v2's.
I plan to submit Purrence to the Open SLM Leaderboard shortly. Its score sits between Pulvis v2 and Pulvis v1 and would place it third among the published models under 3M parameters in the snapshot from 5 October 2026.
In the 5 October 2026 snapshot, the local Purrence score of 8.4873 falls between Pulvis v2 and Pulvis v1. The blog describes the comparison and submission status.
Pausing here
I need to focus on university applications and school, so I won't be able to work on this project for the foreseeable future. This is where I'm leaving Purrence for now.