Arthur B.
Arthur B. Arthur Breitman. Machine learning, functional programming, applied cryptography, and these days mostly #tezos. Husband of @breitwoman, oligocoiner.

A Piano in 900 Characters

my work, AI-generated overview

Spectrogram of the golfed piano playing five chords, quiet to loud

I wanted to see how small a convincing piano could be. Here is the whole voice:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
from numpy import*
def synth(m,v=64,d=3):
 e=m-60
 a,s,D,k=[interp(v,[36,75,101],y)for y in([2.43,2.19,1.84],[.04,.088,.094],[3,3,2],[.02,.035,.06])]
 F=440*2**((m-69)/12)
 n=arange(1,65)
 f=n*F*(1+exp(-7.79+.0653*e)*n*n)**.5
 j=f<=12e3
 n,f=n[j],f[j]
 A=n**-clip(a+s*e,1.5,3.6)/(1+(f/6e3)**4)**.5
 A/=A.max()
 T=maximum(.835*exp(-.0277*e)*(F/f)**.4,.05)
 N=int(d*44100)
 t=arange(N)/44100
 g=random.default_rng(m)
 w=fft.rfftfreq(N,1/44100)
 E=A[:,None]*(.85*exp(-t/T[:,None])+.15*exp(-8*t/T[:,None]))
 S=1 if m<=28 else 2 if m<=40 else 3
 o=zeros(N)
 for c in[0]if S<2 else[(2*i/(S-1)-1)*D for i in range(S)]:
  o+=(E*sin(2*pi*(f*2**(c/1200))[:,None]*t+g.uniform(0,2*pi,len(f))[:,None])).sum(0)
 o=o/S*(1-exp(-t/.004))
 z=fft.irfft(fft.rfft(g.normal(0,1,N))/(1+(w/1500)**2)**.5,N)
 o+=k*z/(abs(z).max()+1e-9)*exp(-t/.006)
 o=fft.irfft(fft.rfft(o)*10**(polyval([-.564,-.814,1.385],log2(w/1e3+.02))/20),N)
 return float32(o)

synth(midi, velocity, seconds) returns 44.1 kHz audio. There are no samples, only a sum of sine waves, and every constant in it was measured from the University of Iowa piano recordings at pp, mf and ff.


You can tell it isn’t a real piano. For its complexity, though, it’s surprisingly good. The upper range is the weakest part. Here it is playing Chopin’s Nocturne Op. 9 No. 2:

What the lines do

  • f=n*F*(1+B*n*n)**.5: stiff strings are inharmonic, so their partials run progressively sharp. log B is linear in pitch.
  • A=n**-slope: partial amplitudes fall off as a power of the partial number. The slope depends on pitch and velocity, which makes the bass rich, the treble pure, and loud notes brighter. Velocity is continuous: interp blends between the three measured dynamics.
  • T and E: each partial decays at its own rate, faster for higher notes and higher partials. The decay has two stages, a quick initial drop and then a long tail.
  • S and D: one, two or three strings depending on register, detuned by a few cents. That produces the slow beating you hear in a real piano.
  • z: the hammer knock, a 6 ms burst of low-passed noise.
  • The last line: a fixed soundboard EQ, quadratic in log-frequency. I started with a nine-point lookup table and swapped in a fitted quadratic to save characters.

The golfed function is bit-identical (max error about 1e-13) to the fitted model it came from. I also wrote a C port, about 1500 characters with only math.h and no FFT, and a Julia port. The Julia version came out longer than the numpy one, mostly because numpy has interp built in.

Measure, don’t fit

The constants were the hard part. I first fit everything by gradient descent through a spectrogram loss, and it was a disaster. Inharmonicity went negative, the decays collapsed, and brightness came out inverted. A magnitude-spectrogram loss barely constrains any of those quantities.

What worked was a lock-in analysis of every recorded note: heterodyne each expected partial down to DC and low-pass it. That gives each partial’s exact frequency and amplitude envelope directly. B, the decay laws and the unison detune were read off those measurements and pinned. Only the excitation shape, the one thing the audio loss can actually see, was left to the optimizer.

The oboe

The oboe isn’t too bad either. Here it is with the piano accompanying:

It needs a different model. A reed is driven, not struck, so the tone sustains and the partials are almost perfectly harmonic. Its character comes from fixed resonances in the bore rather than from how it decays. I fit a formant envelope to 35 Iowa oboe notes: a +28 dB peak at 1.29 kHz and a second one near 3 kHz, on a −10 dB/octave tilt. On top of that sit vibrato and some breath noise.

The trick that makes it sound alive is that the formant stays fixed in absolute frequency. As the vibrato moves the pitch, each partial slides up and down the formant curve and its loudness shimmers, the way it does on a real oboe.

This was part of a larger project on Sethares-style spectrum tuning: nudging partials so that equal-tempered intervals become dissonance minima. The piano’s spectrum is too steep for that to work, about 3 of 12 intervals. The oboe’s brighter spectrum gets 9 of 12, tritone included.

The code is on GitHub at murbard/golfed-synths. It’s numpy only, and python demo.py renders the clips above.

comments powered by Disqus