aldus.nexus

Neural network

A small network learns live, drawn as a constellation, with errors flowing backwards through it. Place stars on the plane and watch it carve out a nebula for each colour.

positive weight negative weight forward pulse gradient
epoch 0 click the plane to

Step-through: one star, forward and backward

The numbers

losslog scale
accuracy
gradient size per layermean |dL/dw|, log

Draw a digit

what it sees: centred and shrunk to 16 × 16
its guess
confusion matrix, unseen test digits

No downloaded dataset: when the page opens it draws a few thousand digits itself, typed in the site's fonts and this device's fonts or drawn as pen strokes, each turned, slanted, stretched, shifted and given a different stroke width, then trains a 256-64-10 network on them in a background worker. The weights are kept in this browser, so the next visit skips straight to guessing.

How it works

Each neuron takes the numbers from the layer before, multiplies each by a weight, adds a bias and squashes the total through an activation such as tanh. Two inputs (a star's x and y) flow through a few hidden layers to one output per colour, and softmax turns those into probabilities.

Training compares the answer with the star's true colour using cross-entropy loss. Backpropagation then walks backwards, using the chain rule to work out how much each weight was to blame, and gradient descent nudges every weight a little the other way. Momentum keeps the nudges rolling in a consistent direction.

The nebula is the network's answer at every point of the plane. Hollow rings are test stars it never trains on, so test accuracy tells you whether it has learned the shape or just memorised the stars. Stack many sigmoid layers and watch the gradient chart: the early layers barely learn, because every factor of da/dz is at most 0.25.

The draw-a-digit reader is the same machine, bigger: 256 inputs (one per cell of the 16 × 16 pad), 64 ReLU neurons and 10 outputs. Softmax turns the 10 output scores into probabilities that add up to 1, and training lowers the cross-entropy, minus the log of the probability given to the right digit. Its gradient at the outputs is simply p minus the one-hot answer. The confusion matrix counts, for test digits it never trained on, which digit it guessed against which was true: the diagonal is right, everything else is a mix-up.

nothing in the network knows what a spiral is; the shape emerges from millions of tiny local nudges, each one just the chain rule.

∂L/∂w = ∂L/∂a × ∂a/∂z × ∂z/∂wz = Σ w·a + b is a neuron's weighted sum, a = f(z) its output, L the loss. Each factor is local, so gradients are multiplied backwards layer by layer.
pₖ = e^zₖ / Σⱼ e^zⱼsoftmax: scores z become probabilities p that sum to 1
L = -log py, ∂L/∂zₖ = pₖ - [k = y]cross-entropy for the true class y, and its neat gradient
w ← w - η · ∂L/∂wη is the learning rate. Too small and it crawls; too big and it overshoots.

Challenges