aldus.nexus

Tuner and spectrogram

Every sound as a scrolling map of frequencies, and the pitch found the way a good tuner finds it: YIN compares the wave with delayed copies of itself, and the delay that matches best is one period. Play the test tone, music or your own instrument.

Microphone: only after you press it, and the browser asks first. The sound is analysed on this device and never recorded or sent anywhere. The mic switches off when you press stop or pause, hide this tab or leave the page. Files stay in your browser too.

Test tone a quiet, soft tone: click a key to play it, Stop or T to silence it

C3 to B4

The numbers

step 1: difference d(τ)dips where the wave repeats
steps 2 to 4: normalised d′(τ)first dip under θ wins
pitch over timeHz, last 12 s
notes heardtime on each note name

How it works

Two ways to look at the same sound. The spectrogram asks which frequencies are present; the tuner asks how long the wave takes to repeat. Each step below is filled in live from …. Related pages: the Visualiser uses the same music input and FFT to drive its scenes, Fourier epicycles shows the FFT's butterfly and turns shapes into frequencies, Oscilloscope music goes the other way and draws with sound, and Euclidean rhythms covers the other half of music, time. The visualiser also has this spectrogram and YIN pitch line as a scene: open it there. The How it works guide has the short version.

1Samples to frequencies

The sound arrives as fs numbers a second. The browser's analyser takes the latest N of them, multiplies by a window that fades the ends, and runs a fast Fourier transform, which splits them into N/2 frequency bins an equal distance apart. More samples give finer bins but a blurrier sense of time, so a spectrogram always trades one for the other.

Δf = fs / Nbin spacing… / … = …

X_k = Σₙ wₙ xₙ e^(−2πikn/N), f_k = k · Δfbin k measures frequency f_kwindow length …

2A map of the sound

Each analyser frame becomes one thin column: loudness in decibels picks a colour from deep indigo to white-gold, and the columns scroll left. Pitch is drawn on a log axis, so every octave gets the same height, as on a piano. That squeezes the high bins together and stretches the low ones: at the bottom one bin covers several pixel rows, and the picture is interpolated.

L = 20 · log₁₀ |X_k| dBcolour from L between the floor and −28 dBfloor …

f(y) = f_lo · (f_hi / f_lo)^(1 − y/H)row y on a log axis…

3YIN step 1: compare with a delayed copy

Slide a copy of the wave along by τ samples and add up the squared differences. If the sound repeats every T samples, the copy lines up again at τ = T and the sum dives towards zero. The first chart shows d(τ) live; its dips sit at the period and its multiples.

Done literally that is W × τ_max multiplications, about 2.5 million a frame. Expanding the square leaves two running sums of squares and one correlation r(τ), and the correlation of every lag at once comes from a single FFT, about ten times faster here.

d(τ) = Σ_{j=0}^{W−1} (x_j − x_{j+τ})²W samples, every lag τ up to τ_maxW = …, τ_max = … (lowest note …)

d(τ) = m₀ + m_τ − 2 r(τ), r = IFFT(conj(A) · B)m_τ = Σ_{j=τ}^{τ+W−1} x_j², from prefix sumsthis frame: …

4Step 2: normalise

d(τ) is zero at τ = 0 and wanders with loudness, so it can't be compared with a fixed bar. Dividing by its running average fixes both: d′ starts at 1, stays near 1 where nothing repeats and dips towards 0 at a period, whatever the volume.

d′(τ) = d(τ) / ((1/τ) Σ_{j=1}^{τ} d(j)), d′(0) = 1cumulative mean normalised differenceat the chosen lag d′ = …, clarity 1 − d′ = …

5Step 3: the first good dip

Take the first lag where d′ drops under the threshold θ and follow it down to the bottom of that dip. Taking the first, not the deepest, avoids picking twice the period, an octave too low. Set θ too high and a dip at a half period sneaks in, an octave too high; set it too low and nothing qualifies, so the deepest dip anywhere is used instead.

τ* = min { τ : d′(τ) < θ }, then down to the local minimumelse τ* = argmin d′(τ)θ = …: …

6Step 4: between the samples

The true period rarely lands on a whole sample: at 48 000 samples a second an A4 repeats every 109.09 of them. A parabola through the dip and its two neighbours finds the bottom between samples, which matters at high notes where one sample is many cents.

τ′ = τ* + (d′₋ − d′₊) / (2(d′₋ − 2d′₀ + d′₊))d′₋, d′₀, d′₊ at τ* − 1, τ*, τ* + 1τ* = … → τ′ = …

f₀ = fs / τ′the pitch…

7Note names and cents

In equal temperament each semitone multiplies the frequency by 2^(1/12), so pitch is a logarithm of frequency. The nearest note comes from rounding, and the needle shows the rest in cents, hundredths of a semitone. Most ears notice about 5 to 10 cents; the green zone is ±5.

n = 69 + 12 · log₂(f / A4)MIDI note number, A4 = 69…

cents = 1200 · log₂(f / f_ref)f_ref = A4 · 2^((n − 69)/12)…

8Why not just take the loudest peak?

Many instruments are louder in a harmonic than in the fundamental, and some sounds have no fundamental at all: the missing-fundamental test tone plays only harmonics 2 to 7, yet you still hear the low note, and so does YIN, because the wave still repeats at the low period. The loudest FFT bin lands on a harmonic instead, and its bins are Δf apart, which is many cents wide at the bottom.

f_peak = Δf · argmaxₖ |X_k|naive guesspeak … against YIN …

1 bin at f = 1200 · log₂((f + Δf) / f) centshow coarse the FFT is…

One subtraction, repeated at every delay, hears the pitch your ear hears, even when the note itself is missing from the sound.