Day 55 · 2026.08.16

Mathematics & Music, Deeper

Scales are number theory, rhythm is group theory, hearing is harmonic analysis, composition is long-range dependence
"Music is the pleasure the human soul experiences from counting without being aware that it is counting." — Leibniz, letter to Goldbach, 1712

Scales as Diophantine Approximation

Why It Had to Be 12
Number Theory
Intuition First

Music has only two near-natural intervals: the octave (frequency ×2) and the perfect fifth (×3/2). The trouble is that they are incompatible — walk upward by fifths and you never return to where you started: $(3/2)^m = 2^n$ would require $3^m = 2^{n+m}$, which unique factorization vetoes outright.

So a scale must be a compromise, and compromises have optima: cut the octave into $q$ equal parts, and ask which $q$ lands the perfect fifth closest to a grid point. The musical question has just become a number-theoretic one — approximating an irrational by a rational with a small denominator.

Formal Definition

Let $x=\log_2(3/2)=0.5849625\ldots$, the fraction of an octave occupied by a perfect fifth. We seek $p/q$ minimizing $|x-p/q|$, where $q$ is the number of equal divisions and $p$ the number of steps in a fifth. The continued fraction and its convergents:

$x = [0;1,1,2,2,3,1,5,\ldots] \Rightarrow \tfrac{1}{2},\ \tfrac{3}{5},\ \tfrac{7}{12},\ \tfrac{24}{41},\ \tfrac{31}{53}$

The best-approximation theorem guarantees that convergents satisfy $|x-p/q| < 1/q^2$, and that no fraction with denominator at most $q$ comes closer. Here $3/5$ is the pentatonic scale and $7/12$ is twelve-tone equal temperament (a fifth spans 7 semitones). Taking 1200 cents to the octave, the pure fifth is 701.96: equal temperament gives 700 (off by −1.96), while 53 divisions give 701.89 (off by −0.07, far below the roughly 5-cent threshold of human discrimination).

0 cents +20 −20 5 +18.0 7 −16.2 12 −1.96 19 31 41 53 −0.07 error of the perfect fifth in N-tone equal division (cents)
Why It Is Beautiful

The pentatonic and the twelve-tone scale are not cultural accidents; they are two adjacent terms of the same continued fraction. The Chinese sān-fēn-sǔn-yì cycle, the Pythagorean chain of fifths, the Arabic 17-tone system, the 22 śrutis of India — civilizations working in isolation all landed near this sequence of convergents. Number-theoretic convergence pushed different cultures onto the same grid points.

More telling still is 53: mathematically near-perfect (Nicholas Mercator and the Chinese 60-lü tradition both brushed against it), yet adopted by no mainstream instrument. Between the optimal solution and the playable one stands a wall of engineering — what twelve-tone equal temperament won on was never accuracy, but the number of fingers and the width of a key.

Applications

The same continued-fraction machine runs elsewhere: orbital resonance (Io, Europa and Ganymede locked at 1:2:4), gear-tooth design, the Metonic cycle behind calendar intercalation, divider ratios in phase-locked loops. KAM theory supplies the mirror image: the frequency ratio hardest to approximate by rationals (the golden ratio) marks the orbit least likely to be torn apart by resonance — the gaps in Saturn's rings and the tuning of a piano pose the same question.

Essence + A Question
A scale is where one act of Diophantine approximation happened to land: 12 is 12 because $7/12$ is the best small-denominator rational approximation to $\log_2(3/2)$.
Question: if the human ear were an order of magnitude sharper, would pianos have grown 53 keys — or does the body's constraint always outrank the mathematics?

Rhythm as Group Theory

Subsets of a Cyclic Group
Group Theory · Combinatorics
Intuition First

Arrange the $n$ pulses of a bar in a circle and blacken the struck positions — a rhythm is simply a subset of that ring. Entering the same drum pattern on a different beat is just "another door into the same rhythm": mathematically, the rotation action of the cyclic group $\mathbb{Z}_n$, with one rhythm = one orbit (a necklace). "How many 8-pulse rhythms are there?" thereby becomes a counting problem.

A second intuition: good drum patterns tend to spread $k$ onsets as evenly as possible across $n$ slots. When $k$ does not divide $n$, "as evenly as possible" needs an algorithm — and that algorithm turns out to be Euclid's.

Formal Definition

A rhythm is a subset $A\subseteq\mathbb{Z}_n$ under the rotation $A\mapsto A+1$. By Burnside's lemma, the number of rhythm necklaces of length $n$ is:

$\displaystyle \frac{1}{n}\sum_{d\mid n}\varphi(d)\,2^{n/d}$

Here $\varphi$ is Euler's totient and $2^{n/d}$ counts the patterns fixed by a rotation of period $d$. For $n=8$ the answer is 36 — 256 binary patterns collapse into 36 genuinely distinct rhythms. The Euclidean rhythm $E(k,n)$ comes from Bjorklund's algorithm: run the Euclidean algorithm on $(k,\,n-k)$ and you obtain a maximally even set in $\mathbb{Z}_n$. $E(3,8)=\texttt{x..x..x.}$ is the Cuban tresillo; $E(5,16)$ is the bossa nova clave.

E(3,8) tresillo · gaps 3-3-2 E(5,16) bossa nova · gaps 3-3-4-3-3
Why It Is Beautiful

The most startling identity is $E(7,12)$: spread 7 points maximally evenly over 12 slots and you get the major scale. One and the same "maximal evenness" algorithm yields the world's drum patterns along the time axis and the Western diatonic scale along the frequency axis — pitch and rhythm turn out to be two slices of a single combinatorial object.

And a cold piece of corroborating evidence: Bjorklund's algorithm was written to schedule timing pulses for the spallation neutron source at Los Alamos. It has nothing to do with music, yet it computes the same traditional rhythms — Toussaint checked dozens of world rhythms and found almost all of them are Euclidean.

Applications

"Spread $k$ events as evenly as possible over $n$ slots" is daily bread in distributed systems: timer division, CPU quota scheduling, token issuance in rate limiters, virtual-node placement in consistent hashing — and Bresenham's line-drawing algorithm is literally the same procedure. The group-theoretic half supplies a rhythm fingerprint: compare orbits rather than sequences and identification of looped samples becomes start-point invariant for free.

Essence + A Question
A rhythm is a subset of a cyclic group, and the compelling ones are the maximally even ones — where "maximally even" is computed by Euclid's algorithm.
Question: a clave really does sound different when entered on another beat, so does the equivalence class "orbit" erase information the ear genuinely hears? When should mathematical equivalence and perceptual equivalence part ways?

Time–Frequency Uncertainty and Constant Q

Music's Natural Coordinate Is Logarithmic
Harmonic Analysis
Intuition First

To decide "which note is this?" you must listen for a while — pitch is read off by counting cycles, and the shorter the window the fewer cycles there are, so the pitch blurs. Conversely, to know exactly "when was it struck?" you need an extremely short window, and then there is almost no frequency information left.

This is not a shortcoming of equipment but a property of waves: a tone that is both short and pure does not exist mathematically. The Fourier transform (Day 40) tells you which frequencies are present, not when they occur; ask both questions at once and there is a non-negotiable bill to pay.

Formal Definition
$\Delta t \cdot \Delta f \;\ge\; \dfrac{1}{4\pi}$

$\Delta t$ is the standard deviation of the signal's energy along time, $\Delta f$ its standard deviation along frequency; equality holds only for a Gaussian window — which is exactly why Gabor atoms carry a Gaussian envelope. The engineering version is easier to remember: a window of $T$ seconds buys roughly $1/T$ Hz of frequency resolution.

The constant-Q transform spends the budget differently: it holds the relative resolution fixed, $Q=f_k/\Delta f_k$, with bins spaced logarithmically as $f_k=f_0\,2^{k/b}$. Resolving a semitone demands $\Delta f/f = 2^{1/12}-1\approx 5.9\%$, i.e. $Q\approx 17$: long windows low, short windows high.

uniform STFT: equal tiles constant Q: logarithmic tiles time → frequency ↑ (linear) sharp in time up high, sharp in frequency down low highlow
Why It Is Beautiful

Fourier hands you a uniform frequency grid, but music lives on a logarithmic one: near 30 Hz a semitone is only 1.8 Hz wide, while at 3000 Hz it is nearly 180 Hz. A uniform STFT therefore cannot separate the bass and squanders resolution in the treble. Constant Q gives each band a window proportional to its wavelength, which amounts to redistributing the uncertainty budget along music's own geometry.

This is more than an engineering trick: logarithmic frequency is not a notational convention — the cochlea's place-to-frequency map is already roughly logarithmic, and wavelet analysis independently produces the same family of scale-self-similar atoms. Physiology, mathematics and music theory converge on one coordinate system. It even hard-codes a rule of composition: low notes need longer to be identified, so a bass is natively unsuited to very fast runs.

Applications

The phase vocoder decouples time-stretching from pitch-shifting, the foundation of every tuning and stretching plug-in; CQT and mel spectrograms are the front end of nearly all music AI (transcription, chord recognition, the input to audio generative models). In quantum mechanics the very same inequality is called Heisenberg's principle — position–momentum and time–frequency are one and the same pair of Fourier conjugates.

Essence + A Question
"Short and precise at once" is mathematically forbidden; every perceptual trade-off in music between speed and register is a corollary of this inequality.
Question: people distinguish onset differences under 10 ms, apparently breaking the resolution wall — does hearing evade uncertainty, or was it never a linear transform to begin with?

Algorithmic Composition and Long-Range Structure

The Gap Between Exponential Memory and Power-Law Form
Probability · Machine Learning
Intuition First

The most naive way to make a machine compose: tabulate what usually follows the previous few notes, then sample. Hiller and Isaacson's Illiac Suite of 1957 — the first computer-composed score — worked exactly this way. The outcome is emblematic: locally plausible, globally formless. Thirty seconds sound like music; three minutes reveal that nothing was being said.

The failure can be stated precisely: a Markov chain's memory decays exponentially with distance, while musical structure is a power law — a second section echoing the first means correlations survive hundreds of notes. No increase in the order $k$ closes that gap.

Formal Definition
$\displaystyle p(x_{1:T}) = \prod_{t=1}^{T} p\!\left(x_t \mid x_{

This is the autoregressive factorization: every note is conditioned on the entire history, $x_{A power law has no characteristic scale; an exponential does — this is a structural gap, not a badly tuned parameter.

Markov: short adjacent arcs only form: distant echoes
Why It Is Beautiful

This translates "music has form" into a measurable statistical signature. A power law means no characteristic scale, and no characteristic scale means self-similarity (the fractals of Day 17): the hierarchy of theme, development and recapitulation leaves the same statistical fingerprint as a critical phenomenon.

The repairs are equally mathematical, and there are only two. Write the hierarchy into the model explicitly (Schenkerian analysis, the tree grammars of GTTM), or let the model see hundreds of steps away directly — attention in a Transformer is precisely "an edge between any two positions", and relative position encoding turns "how many beats since this last appeared" into a learnable quantity. That Music Transformer first produced piano pieces with genuine recapitulation amounts to tearing down the wall of exponential decay.

Applications

Music generation models (over MIDI events or audio tokens) and accompaniment and continuation tools are the direct payoff. The same long-range-dependence analysis is carried over to evaluating language models (why low perplexity still drifts off topic over long documents), detecting long-range correlation in DNA, and diagnosing cross-file dependence in code completion. Xenakis's other road is still open too: model the distribution rather than the notes — today's diffusion models are its distant relatives.

Essence + A Question
The hard part of composing was never the next note; it is making note 200 remember note 1.
Question: given unlimited context, will a sense of form emerge on its own, or must hierarchy be represented explicitly before it can be learned?

Going Deeper

Equal temperament's major third is off by +13.7 cents, seven times the error of the fifth — why is it the more acceptable one?
Because audibility tracks beat rate, not cents. A fifth mistuned by −2 cents beats about 0.9 times a second near middle C, essentially imperceptible; a third off by 13.7 cents beats close to 10 times a second, squarely in the roughest region. It is tolerated for two reasons: piano partials are not exact integer multiples anyway (string stiffness produces inharmonicity), and three centuries of training have re-encoded that slight muddiness as "brightness". A small mathematical error and a large perceptual one are being read off two different rulers.
Scales and rhythms share the single theorem of maximal evenness — so why does an evenly divided rhythm sound mechanical while the major scale sounds natural?
What matters is whether the even set retains a little asymmetry. The gap sequence of $E(7,12)$ is 2-2-1-2-2-2-1, and the inequality of the two step sizes is exactly what creates orientation — you can hear which degree of the scale you are on. A perfectly equal rhythm (or the whole-tone scale) has identical gaps, every rotation returns it to itself, and hearing loses its reference point and therefore its tension. Beauty seems to live at "almost even, with one symmetry broken" — perfect symmetry carries zero information, the same intuition as the symmetry breaking of Day 18.
If "beautiful" can be captured by mutual information or compression rate, does it become an optimizable objective?
Schmidhuber's "compression progress" theory claims exactly that: beauty is the moment an observer's compressor is improving fastest — neither maximally compressible (boring) nor incompressible (noise). The difficulty is that it is observer-relative and moves in time: the tenth hearing of a piece brings no further progress although the compression rate is unchanged. The instant a fixed metric becomes the optimization target, Goodhart's law takes hold — the model learns to manufacture statistical surprise rather than musical surprise.
Is there a single thread running through the four concepts?
Yes: music imposes discrete structure on continuous physical quantities, and the cost at every layer can be computed. Frequency is cut by $\mathbb{Z}_{12}$, at the cost of a Diophantine residual; time is cut by $\mathbb{Z}_n$, at a cost minimized by maximally even sets; hearing is constrained by an inequality even before discretization, at the cost of a time–frequency resolution product; composition must rebuild long-range order over discrete symbol strings, at the cost of a mutual-information decay law. Discretization is always lossy, and musical form grows in the shape of those losses.