Array Languages #5: X_eTaL-ML, Where the Machine Learning Is
2039 words • 11 min read • Abstract

| Resource | Link |
|---|---|
| The repo | softwarewrighter/X_eTaL-ML — machine-learning demos and the ML libraries they are made of |
| Live | the catalog — all ten demos in your browser; start with the Tiny CNN |
| Run and read | run them yourself · the code, cross-referenced |
| The pieces | sw-ml-study/sw-mlpl · Rosetta M · X_eTaL |
| The thread | ML #9 · Made Visible #4 |
| Prior post | Array Languages #4: X_eTaL-extensions |
| Comments | Discord |
Why an array language for ML
The repo’s README makes the case in a paragraph. Most of what an ML framework hides behind layers and modules is array algebra. A dense layer is one inner product and an addition. Softmax is an exponential, a row sum and a division. Attention is two matrix products and a softmax. A mixture-of-experts router is a matrix product and a top-k. In X_eTaL each of those is one short, typed expression over whole arrays, so a demo can show the model itself: the weights, every intermediate array with its shape, and the moment a loop nest turns out to be a single array transformation. And because the types are inferred and checked, a shape mistake in a layer is an error before the forward pass, not a wrong number after it.
What runs today
Ten demos, all live in the browser and all runnable at the command line from a clone. Each page shows all the code it runs, beside the arrays that code computed. Each tile links to its live page:
Tiny CNN
backprop microscope
CNN training
training live
attention microscope
MoE routing
network macro
k-means
embedding explorer
1.58-bit network
They fall into four kinds:
- Watching a model think. The Tiny CNN reads a digit you draw, every 3 by 3 window of the picture taken at once as nine rotated copies of it. The attention microscope shows one head over your sentence, and “tired” finds the animal. The MoE router sends each token to two of 16 experts, and the 1.58-bit network compares one classifier at four precisions, down to weights of -1, 0 and +1.
- Training, in X_eTaL. Three demos train. The backprop microscope shows one step with every array, and checks every gradient by nudging its weight. Training live fits a spiral with Adam while the decision regions bend to follow it. CNN training trains a small CNN from random weights on 600 handwritten digits, in your browser, through softmax, a dense layer, max-pooling, ReLU and the convolution, every gradient checked by finite differences.
- Classic machine learning. k-means clusters five blobs step by step, and shows a poor start settling wrong where a farthest-first start finds all five.
-
Embeddings. The embedding explorer makes word embeddings from counts, in X_eTaL: 4,000 short sentences over 75 words, each word described by 64 numbers from the words that occur near it. PCA turns the 64 dimensions into 3, and in the cloud you can turn, animals, foods, colors, numbers and places gather. The whole embedding table is one line:
E ← ᵉᵐu̲nit 64 ᵉᵐp̲roject ᵉᵐp̲pmi (v, 2) ᵉᵐc̲ooccur ids
The lines that do the work stay short. Every 3 by 3 window of a picture, and a convolution layer’s gradient over a whole batch:
w ← -1 0 1 o̲-₂ -1 0 1 o̲-₂ x
GK ← DM '+ '× i̲nner o̲\ V
Six libraries, and a network in one line
The demos are made of six libraries in the repo. Three are new this week: Quant and Conv were moved out of the demos that first needed them, so the next demo can use them, and Embed arrived with the embedding explorer.
- NN is the vocabulary: activations, softmax by row at any rank, dense layers, argmax, one-hot, loss and accuracy.
- Quant puts weights in fewer bits: FP16 and bfloat16 rounding, symmetric integers, and ternary codes with their scale. The 1.58-bit network now runs on it.
- Conv is convolution for a batch of pictures: 3 by 3 windows, the filters as one matrix product, max-pooling, and their backward passes. The CNN training demo now runs on it.
- Embed makes word embeddings from counts: co-occurrence counted by sorting, PPMI, a fixed projection, and cosine neighbors.
- Learn is classic machine learning, each method a fit and a predict: k-means, k nearest neighbors, PCA and logistic regression.
- Net is a macro library, the second meaning of Extensible from post #3 put to work. One line writes a network:
"net:" u̲se< "Net"
"u:d_eep c" ⁿᵉᵗm̲odel< "2 16 relu 16 relu 3 softmax"
That call expands into the code that loads the weights, checked, and an ordinary forward function of NN calls, and xetal expand shows all of it. A second macro, net:t_rain<, writes the network’s backpropagation and an Adam step. The network macro demo lets you type a spec of your own and train it in the page.
Tuples, new in X_eTaL this week, came in exactly where this repo had asked for them. A training state is several arrays of different shapes: two weight matrices, Adam’s two running averages for each, and a step count. Before tuples it had to be packed into one long vector. Now it is one tuple, taken apart by name, and p̲ower iterates it as one value:
ᵘa̲dam ← { (W1, W2, M1, M2, V1, V2, k) →
t ← 1.0 + k
(G1, G2) ← W1 ᵘg̲rad W2
m1 ← (0.9 × M1) + 0.1 × G1
m2 ← (0.9 × M2) + 0.1 × G2
v1 ← (0.999 × V1) + 0.001 × G1 × G1
v2 ← (0.999 × V2) + 0.001 × G2 × G2
(W1 − lr × (m1, v1) ᵘm̲ove t, W2 − lr × (m2, v2) ᵘm̲ove t, m1, m2, v1, v2, t)
}
s2 ← 300 'ᵘa̲dam p̲ower s1
TTTML: an older kind of learning
One learner sits outside all this. TTTML, the tic-tac-toe machine from TBT #12, began as an APL workspace written for sw-apl’s 1975 mode, and its X_eTaL port lives in the X_eTaL repo as a demo, not here. It learns a different way: a table with one number per board position, filled in by playing itself, with a position and its rotations and reflections counted as one. There is no network, no gradient and no backpropagation in it. The X_eTaL-ML demos are neural networks trained by gradient descent, the approach behind today’s models. TTTML is worth keeping as the old approach in miniature, and it shows how far a table and self-play go, but it is not where this repo’s machine learning is heading.
The pieces
X_eTaL is not the ML language. It is a precursor: a research language whose job is the notation, and whose results feed a later language that is not ready to write about yet.
The pieces are these. sw-MLPL has the machine-learning built-ins — softmax, sigmoid, relu, Adam, autograd — and takes ASCII, no glyphs required. X_eTaL has the conciseness and the typed, decorated surface, and also takes plain ASCII. Rosetta M has the visual side, the four faces and the callouts, and its notation face does use glyphs: placeholders, drawn by hand for the mock-up. The language after X_eTaL takes ASCII keystrokes and pretty-prints sequences of two or more characters into shorter glyphs, the way X_eTaL already pretty-prints r_ev into r̲ev. The new glyphs for the ML vocabulary are not designed yet, and when they are, each should correspond to one sw-MLPL built-in. So X_eTaL is the notation that could be extended to replace Rosetta M’s placeholders: the conciseness of X_eTaL and the built-ins of sw-MLPL, implementing something visual like Rosetta M. And before the language is specialized for ML, or a new ML-focused language is derived from it, the features that language would need are being tried in X_eTaL as macros rather than as new syntax.
What comes next
The repo’s next stretch is libraries first, then the demos that need them. Quantization, convolution and embeddings are done. Attention moves into a library of its own next, followed by layer norm and sampling. Then a MicroGPT: its forward pass in X_eTaL on weights trained elsewhere, sampling names. Then an optimizer library and the libraries on the live site. A world model and diffusion from noise stay deferred, waiting on training speed and a learned denoiser.
Alongside that plan, and in the X_eTaL language repo itself, work has begun on macros for the kind of math equations ML is written in. The first target is the 2D convolution that ML #9 ended on, a triple sum over input channels and kernel rows and columns. The aim is one equation, written once, giving three things: math notation you would recognize from a paper, the ordinary X_eTaL array code it expands into, which xetal expand can show, and a computation that runs and can be debugged, checked against an independent CNN. The expectation is that macros are a sufficient mechanism to express ML concisely and idiomatically, and the language’s grammar changes only when a change is justified. It is early, a proof of concept still being planned and probed, and this post will be updated when there is something to show.
A few things wait on the language, filed as asks rather than worked around: a grade per row, arrays passed in and out of the browser engine, and e̲ach returning arrays. Speed is measured and guarded, and the CNN’s whole program runs in about a third of a second.
Next in the series
A planned post covers the games, and what a game asks of a language that a demo never does.
Part 5 of the Array Languages series. View all parts
Comments or questions? SW Lab Discord or YouTube @SoftwareWrighter.