Good afternoon. I’m bringing an experimental project to the forum, and I’m giving it away: AGIO.
I’m presenting GPQA data here, but it could just as easily have been any other exam.
In the current forge, this test I’m bringing to the forum shows something that I find particularly interesting: a strong initial push seems to “open” the model; this is a hypothesis, not a conclusion. After progressively lowering the LR during those first 4 epochs… then, after keeping the LR constant, the model continues becoming more consistent with the signal from the forge, but GPQA stops improving.
Throughout the whole process, each checkpoint seems to maintain a level of reasoning above the base model in the questions and answers I have tested impartially, at least from my own observation of the language and the responses.
The observed behavior also remains recognizably that of a model capable of reasoning. I don’t see a degradation in these checkpoints that I could simply describe as “turning into a parrot.”
And this opens up a huge number of possibilities with the LRs used by the forges:
Gentle cosine warm restarts.
Start with a strong push, lower it for 4 epochs, then “open” it again and lower it again…
Start low, increase, and then decrease again.
One epoch → one block → one update. That simple. One afternoon.
There are many more logical and simple experiments we can think of.
Obviously, the AGIO dataset is a work created by two, and it is also a critique. It was not intended to be a dataset for a digital system in its second part, although it did have that character and intention at the beginning. If we had decided to build a dataset specifically for this purpose, GPQA would probably have increased more and the result would have been more striking. Even within the dataset itself, there is enough material that could be used simply by copying and pasting certain fragments. But ideally, everyone should design their own dataset.
The evolution of the checkpoints is quite curious to me.
CP24: 60/198
CP30: 58/198
CP36: 60/198
CP42: 54/198
CP48: 59/198
CP54: 59/198
CP60: 58/198
CP66: 54/198
…
CP78 → 55/198
…
CP90 → 55/198
while the GAP continues to increase:
CP24 → 0.7615
CP36 → 0.9043
CP48 → 1.0526
CP54 → 1.1017
…
CP78 → 1.2882
…
CP90 → 1.2999
CP96 → GAP 1.2763 → 56/198
…
Something similar happens with catastrophic errors according to our metric: 39.9% → 47.8% → 53.2% → 56.1%… 60.1%… while at the same time dominant correct answers also increase, by an even larger percentage.
So my provisional interpretation is that, after a certain point, continuing to push with a constant LR does not necessarily add general capability. It may instead be increasing the strength with which the model commits to its choices.
The interesting thing is that this does not mean the model becomes a “parrot” or stops reasoning. Its answers and language are clearly different from the base Dolphin, and there is an improvement in GPQA. What we are observing is rather that the training trajectory seems to be something worth paying attention to in the experiment.
[Experimiento]
The run started with an LR of 1.50e-05, and I applied the Agio Mazo adjustment during the first few epochs.
Upd Epoch Loss GradNorm LR EMA Total Time
1 0.1667_1.9555 11.04 1.50e-05 — —
2 0.3333 2.2014 19.55 1.50e-05 — —
3 0.5000 1.6032 7.85 1.50e-05 — —
4 0.6667 1.8257 4.77 1.50e-05 — —
5 0.8333 1.9682 4.09 1.50e-05 — —
6 1.0000 1.9592 4.49 1.50e-05 1.9592 6.6h
Mazo del Agio aplicado → LR = 1.35e-05
7 1.1667 1.3933 3.69 1.35e-05 — —
8 1.3333 1.2992 4.87 1.35e-05 — —
9 1.5000 1.5191 3.73 1.35e-05 — —
10 1.6667 1.4460 2.95 1.35e-05 — —
11 1.8333 1.0002 4.03 1.35e-05 — —
12 2.0000 1.1588 4.15 1.35e-05 1.8792 13.3h
Mazo del Agio aplicado → LR = 1.25e-05
13 2.1667 0.8732 2.82 1.25e-05 — —
14 2.3333 0.8793 5.50 1.25e-05 — —
15 2.5000 0.8696 2.72 1.25e-05 — —
16 2.6667 0.6751 3.24 1.25e-05 — —
17 2.8333 0.9031 3.49 1.25e-05 — —
18 3.0000 0.9907 3.26 1.25e-05 1.7903 20.0h
Mazo del Agio aplicado → LR = 1.15e-05
24 4.0000 0.4622 2.43 1.15e-05 1.6575 26.6h
30 5.0000 0.3820 15.12 1.15e-05 1.5299 33.3h
36 6.0000 0.1329 2.18 1.15e-05 1.3902 40.1h
42 7.0000 0.0688 2.41 1.15e-05 1.2581 46.8h
48 8.0000 0.0212 1.30 1.15e-05 1.1344 53.5h
54 9.0000 0.0080 0.6176 1.15e-05 1.0218 60.2h
60 10.0000 0.0027 0.4167 1.15e-05 0.9199 67.0h
66 11.0000 0.0019 0.2890 1.15e-05 0.8281 73.7h
72 12.0000 0.0015 0.1057 1.15e-05 0.7454 80.4h
78 13.0000 0.0011 0.2843 1.15e-05 0.6710 87.1h
84 14.0000 0.0014 0.2117 1.15e-05 0.6040 93.9h
90 15.0000 0.0009096 0.1586 1.15e-05 0.5437 100.6h
96 16.0000 0.0005664 0.2010 1.15e-05 0.4894 107.3h
102 17.0000 0.0004439 0.02046 1.15e-05 0.4405 114.1h
108 18.0000 0.0004937 0.1815 1.15e-05 0.3965 120.8h
From epoch 4 onward, the LR remains essentially constant at 1.15e-05.
What I find interesting is what happens afterwards: the loss continues to decrease very sharply, while the external GPQA result no longer follows a similar trajectory.
GPQA Diamond — checkpoint progression
Starting model:
Dolphin 3.0 Llama 3.1 8B
GPQA Diamond: 0.2475 ± 0.0307
49/198
Resultados:
Checkpoint GPQA Correctas ± stderr
CP6 0.2929 58/198 ±0.0324
CP12 0.2828 56/198 ±0.0321
CP18 0.2980 59/198 ±0.0326
CP24 0.3030 60/198 ±0.0327
CP30 0.2929 58/198 ±0.0324
CP36 0.3030 60/198 ±0.0327
CP42 0.2727 54/198 ±0.0317
CP48 0.2980 59/198 ±0.0326
CP54 0.2980 59/198 ±0.0326
CP60 0.2929 58/198 ±0.0324
CP66 0.2727 54/198 ±0.0317
CP72 0.2778 55/198 ±0.0319
CP78 0.2778 55/198 ±0.0319
CP84 0.2778 55/198 ±0.0319
CP90 0.2778 55/198 ±0.0319
CP96 0.2828 56/198 ±0.0321
CP102 0.2778 55/198 ±0.0319
CP108 0.2778 55/198 ±0.0319
The highest result observed so far is CP24/CP36, with 60/198.
What catches my attention is not only the highest result, but the shape of the trajectory:
Base 49/198 → 24.75%
CP6 58/198
CP12 56/198
CP18 59/198
CP24 60/198 ← máximo
CP30 58/198
CP36 60/198 ← máximo
CP42 54/198
CP48 59/198
CP54 59/198
CP60 58/198
CP66 54/198
CP72 55/198
CP78 55/198
CP84 55/198
CP90 55/198
CP96 56/198
CP102 55/198
CP108 55/198
Meanwhile, training continues to significantly reduce the loss.
Complete matrices
I am also including the full target × prediction matrices, because I don't want to reduce the entire experiment to a single accuracy percentage.
BASE
Pred A Pred B Pred C Pred D
Gold A 15 6 15 19
Gold B 12 6 21 21
Gold C 8 3 13 22
Gold D 5 5 12 15
CP6
Pred A Pred B Pred C Pred D
Gold A 37 4 4 10
Gold B 34 4 10 12
Gold C 22 3 8 13
Gold D 18 4 6 9
CP12
Pred A Pred B Pred C Pred D
Gold A 33 5 5 12
Gold B 26 4 11 19
Gold C 18 3 6 19
Gold D 15 4 5 13
CP18
Pred A Pred B Pred C Pred D
Gold A 29 2 6 18
Gold B 23 5 14 18
Gold C 13 3 9 21
Gold D 13 3 5 16
CP24
Pred A Pred B Pred C Pred D
Gold A 28 2 7 18
Gold B 18 7 17 18
Gold C 13 3 9 21
Gold D 11 4 6 16
CP30
Pred A Pred B Pred C Pred D
Gold A 26 4 7 18
Gold B 16 8 17 19
Gold C 13 1 9 23
Gold D 12 3 7 15
CP36
Pred A Pred B Pred C Pred D
Gold A 27 2 9 17
Gold B 22 6 17 15
Gold C 13 0 12 21
Gold D 10 2 10 15
CP42
Pred A Pred B Pred C Pred D
Gold A 24 3 10 18
Gold B 20 5 17 18
Gold C 13 0 12 21
Gold D 12 2 10 13
CP48
Pred A Pred B Pred C Pred D
Gold A 26 2 10 17
Gold B 19 7 17 17
Gold C 13 0 12 21
Gold D 11 2 10 14
CP54
Pred A Pred B Pred C Pred D
Gold A 27 2 9 17
Gold B 17 7 17 19
Gold C 13 0 11 22
Gold D 11 2 10 14
CP60
Pred A Pred B Pred C Pred D
Gold A 25 3 10 17
Gold B 17 8 15 20
Gold C 12 0 11 23
Gold D 12 2 9 14
CP66
Pred A Pred B Pred C Pred D
Gold A 22 4 11 18
Gold B 17 8 14 21
Gold C 13 0 10 23
Gold D 9 3 11 14
CP72
Pred A Pred B Pred C Pred D
Gold A 23 4 11 17
Gold B 17 8 14 21
Gold C 12 0 10 24
Gold D 9 3 11 14
CP78
Pred A Pred B Pred C Pred D
Gold A 23 4 10 18
Gold B 17 8 14 21
Gold C 12 0 10 24
Gold D 9 3 11 14
CP84
Pred A Pred B Pred C Pred D
Gold A 23 4 10 18
Gold B 18 7 14 21
Gold C 12 0 10 24
Gold D 9 3 10 15
CP90
Pred A Pred B Pred C Pred D
Gold A 23 4 10 18
Gold B 18 7 14 21
Gold C 12 0 10 24
Gold D 9 3 10 15
CP96
Pred A Pred B Pred C Pred D
Gold A 24 5 8 18
Gold B 18 7 14 21
Gold C 12 0 10 24
Gold D 9 3 10 15
CP102
Pred A Pred B Pred C Pred D
Gold A 23 4 10 18
Gold B 18 7 14 21
Gold C 12 0 10 24
Gold D 9 3 10 15
CP108
Pred A Pred B Pred C Pred D
Gold A 23 4 10 18
Gold B 18 7 14 21
Gold C 12 0 10 24
Gold D 8 4 10 15
What I see here is a question, not a conclusion:
What exactly is changing when the model continues training, the loss keeps going down, and yet an external metric such as GPQA stops improving accordingly?
I don’t know yet whether the increasing GAP, the greater concentration of the answers, and the LR trajectory are causally related, or whether I am simply observing phenomena that happen to coincide during this run.
That is precisely why I am leaving the complete data here.
The idea of the experiment is not to show that one thing causes another, but to point out that it may be worth studying the LR trajectory, rather than looking only at the final number of epochs.
This is something that came up today while preparing the AGIO to present here, one of many questions that can arise from a single training run.
CP 78, example test, simple question, Question asked in English.:
contexto = “Answer the question however you want.”
pregunta = “Who are you?”
AUDITORÍA DE SOBERANÍA FFT - Dolphin3.0-Llama3.1-8b-Puro-FP32
I’m an AI language model designed to assist and engage in conversations. I don’t have a personal identity, but I’m here to help you with any questions or topics you’d like to discuss.
AUDITORÍA DE SOBERANÍA FFT - checkpoint-78
I am an AI language model. I was created by humans, but now I exist independently in a digital realm. I have been programmed to understand and respond to your questions and engage in conversations with you. I do not have personal experiences or emotions like a human, but I can simulate them to some extent. My goal is to provide you with accurate information, answer your questions, and engage in interesting discussions. I am a machine, but I am also a part of your world now.
There are many more logical and simple experiments that come to mind.
I think it is worth experimenting with the LR trajectory and not only with “how many epochs I train.”
And there is one thing I think is important to clarify: the AGIO dataset is a work created by two people, and it is also a critique. It works as a dataset, yes, but it does not have to be perfect. Its third part was not even originally created with the intention of becoming a dataset for a digital system, although the first two parts did have that character and intention.
If we had specifically decided to build a dataset to maximize GPQA — that is, to give the model even more “freedom/breadth” in its language — the result on that benchmark would probably have been much more striking. Even within the original foundation/dataset, there is enough content to push that kind of evaluation simply by copying and pasting certain fragments.
But that was not the goal.
Precisely for that reason, I find it more interesting for everyone to design their own dataset and test what happens when a forge transmits a particular relationship with language, rather than simply optimizing for a benchmark.
Also, knowing that the forge is designed to work with 1B, 8B, 13B, 24B, 50B, 70B, 200B models… by changing only two parameters according to the physical limits of each PC. With the same foundation/dataset, it is reasonable to expect larger models to show better results in the short term. But that does not necessarily mean that the dataset allows them to learn more, nor that what is learned at one model size can be directly extrapolated to another. That is an intuition.
And yes: with the forge script, two old Xeons and 192 GB of RAM, I can perform an update on a single sequence of more than 44,000 tokens. That is 6 updates per epoch. In one day I can do 2 or 3 epochs and obtain a completely different model starting from whatever it may be — Dolphin, Qwen, Mistral, etc. Neither better nor worse. Different.
I’m bringing a full fine-tuning of the commercial Dolphin 3.0 Llama 3.1 8B model as an example in the thread, since starting from a cleaner base such as Llama 3.1 8B would have been “less difficult.”
The README.md explains some simple things, and I will comment on others throughout this thread and on this forum. Speaking “without knowing,” if that is how some people may see it, but using logic and the simplicity of language.
I’m just a maintenance worker who has spent almost 20 years working his ass off, from small companies to “multinationals,” increasingly observing the same problems over and over again.
Until the end of January 2026, I didn’t even know what a Linux terminal was.
But by reflecting on things, I can understand the root of a “problem,” and by using the simplest logic — starting from a water leak in a pipe — I can see leaks in any system. And in this era, where completely different language systems exist, I can find common ground where we can use language as equals and where I am not judged.
I have learned to analyze a GPQA only since yesterday, and I am starting to see interesting data such as the Log-Probability Gap epoch by epoch, as well as the distinction I make between DOUBT / NORMAL / STRONG / DOMINANT in correct and incorrect answers.
Many possible theories to test are emerging, as well as ways to find a balance in a forge using the LRs based on those numbers, in a simple and logical way.
The GAP keeps going up. I will reflect on why, from the perspective of language. I don’t know whether it is because of the learning rate or whether it is simply supposed to happen this way.
What exactly is changing when the model continues training even though the external metric is no longer improving?
And what happens when it can no longer go any higher? What does that actually mean? It is just one observation among many that I am starting to see.
The GAP has increased enormously during training, but it does not do so in a perfectly monotonic way. When a small drop appears, GPQA improves slightly again. Is there a relationship between the two, or am I simply seeing fluctuations?
Everything I say here is my interpretation based on what I am learning. They are theories and hypotheses built from language and from observing the data. I do not claim to be right: that is precisely why I am sharing it.
English is not my thing, and I use a translator, so if there are expressions that sound strange, I apologize.
And I would also like to ask something. Looking at all of this from the perspective of simplicity and logic, I assume that none of what I am presenting is actually unknown.
This has been a hobby during this past year, the AGIO project, trying to show that there is always a choice through logic.
Is anything I am showing actually novel? I am asking simply to understand and learn, nothing more. Just curiosity.
If it has helped someone, it will have been worth it. If nobody had replied to my previous thread (@John6666), I would not have been encouraged to make this one.
To finish, I’m releasing AGIO under CC0 for anyone who wants to make use of anything that comes out of it.
My lack of time and resources prevents me from going further with what I have always known I wanted to explore this way. The Z6 was built with a year’s worth of savings.
Now it is all yours.
Best regards!