The architecture follows the standard pipeline: tokenization → embeddings → causal Transformer blocks → LM head → softmax → loss. Because there's no tensor API hiding the logic, you can actually see the weight matrices changing in the logs.
The "No-Magic" Implementation #
Most frameworks treat the computation graph as a black box. In this setup, every number is a Value
object that tracks its own gradient and children:
class Value {
constructor(data, children = [], backward = () => {}) {
this.data = data;
this.grad = 0;
this.children = children;
this._backward = backward;
}
}
When a multiplication happens (y = a × b), the operation stores the local derivatives (dy/da = b, dy/db = a). The backward()
function then performs a topological sort of the graph to apply the chain rule from the loss back to the weights.
Even the neuron implementation is explicit rather than being a matrix operation:
forward(input) {
let output = sum(
input.map((value, i) => value.mul(this.weights[i]))
);
if (this.useBias) output = output.add(this.bias);
if (this.activation === 'relu') return output.relu();
return output;
}
Performance and Results #
Since this avoids optimized BLAS libraries, it's significantly slower than a production LLM agent, but the educational value is huge. The model starts with random weights and fails basic logic tests:
> can human read ?
model: ? ...
expected: human can read. [WRONG]
After going through pre-training and SFT (Supervised Fine-Tuning), it hits a stable criterion where target tokens reach >95% probability, resulting in:
> can human read ?
model: human can read. [CORRECT]
> can cat read ?
model: cat cannot read. [CORRECT]
Technical Breakdown #
Architecture: Two causal Transformer blocks with multi-head self-attention.Optimization: Adam optimizer implemented from scratch.Embeddings: Combined token and position embeddings to handle sequence order.Requirement: Node.js 18.19+
For anyone wanting a deep dive into the actual mechanics of a Transformer without the overhead of a massive framework, this is a perfect practical tutorial.
https://github.com/sekretov/tiny-language-model-neuro-js
Next AI from Scratch: What It Actually Is →