cd /news/machine-learning/array-languages-5-x-etal-ml-where-th… · home › topics › machine-learning › article
[ARTICLE · art-149291] src=blog.softwarewrighter.com ↗ pub= topic=machine-learning verified=true sentiment=↑ positive

Array Languages #5: X_eTaL-ML, Where the Machine Learning Is

Software developer Michael Wright released X_eTaL-ML, a GitHub repository of ten machine-learning demos written in the X_eTaL array language and runnable in a browser or from a command-line clone. The demos include a Tiny CNN that classifies drawn digits, an attention microscope, a mixture-of-experts router that sends each token to two of 16 experts, a 1.58-bit ternary network compared at four precisions, three training demos with gradients checked by finite differences, k-means clustering, and an embedding explorer built from 4,000 short sentences over 75 words with 64-dimensional word vectors reduced to 3 by PCA. The project's README argues that dense layers, softmax, attention and MoE routing are each a single short typed array expression in X_eTaL, so shape mistakes are caught before the forward pass.

read8 min views3 publishedOct 11, 2026

2039 words • 11 min read • Abstract

Resource Link
The repo softwarewrighter/X_eTaL-ML — machine-learning demos and the ML libraries they are made of
Live the catalog — all ten demos in your browser;start with the Tiny CNN
Run and read run them yourself ·the code, cross-referenced
The pieces sw-ml-study/sw-mlpl ·Rosetta M ·X_eTaL
The thread ML #9 ·Made Visible #4
Prior post Array Languages #4: X_eTaL-extensions
Comments Discord

Why an array language for ML #

The repo’s README makes the case in a paragraph. Most of what an ML framework hides behind layers and modules is array algebra. A dense layer is one inner product and an addition. Softmax is an exponential, a row sum and a division. Attention is two matrix products and a softmax. A mixture-of-experts router is a matrix product and a top-k. In X_eTaL each of those is one short, typed expression over whole arrays, so a demo can show the model itself: the weights, every intermediate array with its shape, and the moment a loop nest turns out to be a single array transformation. And because the types are inferred and checked, a shape mistake in a layer is an error before the forward pass, not a wrong number after it.

What runs today #

Ten demos, all live in the browser and all runnable at the command line from a clone. Each page shows all the code it runs, beside the arrays that code computed. Each tile links to its live page:

They fall into four kinds:

  • Watching a model think. TheTiny CNN reads a digit you draw, every 3 by 3 window of the picture taken at once as nine rotated copies of it. Theattention microscope shows one head over your sentence, and “tired” finds the animal. TheMoE router sends each token to two of 16 experts, and the1.58-bit network compares one classifier at four precisions, down to weights of -1, 0 and +1.
  • Training, in X_eTaL. Three demos train. Thebackprop microscope shows one step with every array, and checks every gradient by nudging its weight.Training live fits a spiral with Adam while the decision regions bend to follow it.CNN training trains a small CNN from random weights on 600 handwritten digits, in your browser, through softmax, a dense layer, max-pooling, ReLU and the convolution, every gradient checked by finite differences.
  • Classic machine learning.k-means clusters five blobs step by step, and shows a poor start settling wrong where a farthest-first start finds all five.

Embeddings. Theembedding explorer makes word embeddings from counts, in X_eTaL: 4,000 short sentences over 75 words, each word described by 64 numbers from the words that occur near it. PCA turns the 64 dimensions into 3, and in the cloud you can turn, animals, foods, colors, numbers and places gather. The whole embedding table is one line:

E ← ᵉᵐu̲nit 64 ᵉᵐp̲roject ᵉᵐp̲pmi (v, 2) ᵉᵐc̲ooccur ids

The lines that do the work stay short. Every 3 by 3 window of a picture, and a convolution layer’s gradient over a whole batch:

w ← -1 0 1 o̲-₂ -1 0 1 o̲-₂ x
GK ← DM '+ '× i̲nner o̲\ V

Six libraries, and a network in one line #

The demos are made of six libraries in the repo. Three are new this week: Quant and Conv were moved out of the demos that first needed them, so the next demo can use them, and Embed arrived with the embedding explorer.

  • NN is the vocabulary: activations, softmax by row at any rank, dense layers, argmax, one-hot, loss and accuracy.
  • Quant puts weights in fewer bits: FP16 and bfloat16 rounding, symmetric integers, and ternary codes with their scale. The 1.58-bit network now runs on it.
  • Conv is convolution for a batch of pictures: 3 by 3 windows, the filters as one matrix product, max-pooling, and their backward passes. The CNN training demo now runs on it.
  • Embed makes word embeddings from counts: co-occurrence counted by sorting, PPMI, a fixed projection, and cosine neighbors.
  • Learn is classic machine learning, each method a fit and a predict: k-means, k nearest neighbors, PCA and logistic regression.
  • Net is a macro library, the second meaning ofExtensible frompost #3 put to work. One line writes a network:
"net:" u̲se< "Net"
"u:d_eep c" ⁿᵉᵗm̲odel< "2 16 relu 16 relu 3 softmax"

That call expands into the code that loads the weights, checked, and an ordinary forward function of NN calls, and xetal expand shows all of it. A second macro, net:t_rain<, writes the network’s backpropagation and an Adam step. The network macro demo lets you type a spec of your own and train it in the page.

Tuples, new in X_eTaL this week, came in exactly where this repo had asked for them. A training state is several arrays of different shapes: two weight matrices, Adam’s two running averages for each, and a step count. Before tuples it had to be packed into one long vector. Now it is one tuple, taken apart by name, and p̲ower iterates it as one value:

ᵘa̲dam ← { (W1, W2, M1, M2, V1, V2, k) →
  t ← 1.0 + k
  (G1, G2) ← W1 ᵘg̲rad W2
  m1 ← (0.9 × M1) + 0.1 × G1
  m2 ← (0.9 × M2) + 0.1 × G2
  v1 ← (0.999 × V1) + 0.001 × G1 × G1
  v2 ← (0.999 × V2) + 0.001 × G2 × G2
  (W1 − lr × (m1, v1) ᵘm̲ove t, W2 − lr × (m2, v2) ᵘm̲ove t, m1, m2, v1, v2, t)
}
s2 ← 300 'ᵘa̲dam p̲ower s1

TTTML: an older kind of learning #

One learner sits outside all this. TTTML, the tic-tac-toe machine from TBT #12, began as an APL workspace written for sw-apl’s 1975 mode, and its X_eTaL port lives in the X_eTaL repo as a demo, not here. It learns a different way: a table with one number per board position, filled in by playing itself, with a position and its rotations and reflections counted as one. There is no network, no gradient and no backpropagation in it. The X_eTaL-ML demos are neural networks trained by gradient descent, the approach behind today’s models. TTTML is worth keeping as the old approach in miniature, and it shows how far a table and self-play go, but it is not where this repo’s machine learning is heading.

The pieces #

X_eTaL is not the ML language. It is a precursor: a research language whose job is the notation, and whose results feed a later language that is not ready to write about yet.

The pieces are these. sw-MLPL has the machine-learning built-ins — softmax, sigmoid, relu, Adam, autograd — and takes ASCII, no glyphs required. X_eTaL has the conciseness and the typed, decorated surface, and also takes plain ASCII. Rosetta M has the visual side, the four faces and the callouts, and its notation face does use glyphs: placeholders, drawn by hand for the mock-up. The language after X_eTaL takes ASCII keystrokes and pretty-prints sequences of two or more characters into shorter glyphs, the way X_eTaL already pretty-prints r_ev into r̲ev. The new glyphs for the ML vocabulary are not designed yet, and when they are, each should correspond to one sw-MLPL built-in. So X_eTaL is the notation that could be extended to replace Rosetta M’s placeholders: the conciseness of X_eTaL and the built-ins of sw-MLPL, implementing something visual like Rosetta M. And before the language is specialized for ML, or a new ML-focused language is derived from it, the features that language would need are being tried in X_eTaL as macros rather than as new syntax.

What comes next #

The repo’s next stretch is libraries first, then the demos that need them. Quantization, convolution and embeddings are done. Attention moves into a library of its own next, followed by layer norm and sampling. Then a MicroGPT: its forward pass in X_eTaL on weights trained elsewhere, sampling names. Then an optimizer library and the libraries on the live site. A world model and diffusion from noise stay deferred, waiting on training speed and a learned denoiser.

Alongside that plan, and in the X_eTaL language repo itself, work has begun on macros for the kind of math equations ML is written in. The first target is the 2D convolution that ML #9 ended on, a triple sum over input channels and kernel rows and columns. The aim is one equation, written once, giving three things: math notation you would recognize from a paper, the ordinary X_eTaL array code it expands into, which xetal expand can show, and a computation that runs and can be debugged, checked against an independent CNN. The expectation is that macros are a sufficient mechanism to express ML concisely and idiomatically, and the language’s grammar changes only when a change is justified. It is early, a proof of concept still being planned and probed, and this post will be updated when there is something to show.

A few things wait on the language, filed as asks rather than worked around: a grade per row, arrays passed in and out of the browser engine, and e̲ach returning arrays. Speed is measured and guarded, and the CNN’s whole program runs in about a third of a second.

Next in the series #

A planned post covers the games, and what a game asks of a language that a demo never does.

Part 5 of the Array Languages series. View all parts

Comments or questions? SW Lab Discord or YouTube @SoftwareWrighter.

── more in #machine-learning 4 stories · sorted by recency
── more on @x_etal-ml 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/array-languages-5-x-…] indexed:0 read:8min 2026-10-11 · —