Notes · Speech from first principles

Source–filter model

The claim that a speech sound factorises into three independent stages:

\[P_{\text{out}}(\omega) = U_g(\omega)\,V(\omega)\,R(\omega),\]

a glottal source spectrum, a vocal-tract filter, and a lip-radiation term. Multiplication in frequency means the tract changes only how much of each frequency survives, not which are present: the source picks the harmonic comb and its pitch; the filter shapes the envelope and its formants.

It holds when the system is linear (speech pressures are a fraction of a percent of atmospheric), time-invariant over ~25 ms (the tongue is slow — which is why every front end chops audio into ~25 ms frames), and source and filter don’t interact (the weakest of the three). Granting those, “which vowel?” collapses to “what filter?” — and a recogniser’s whole front end is the machinery for recovering that filter and throwing the source away.

← All notes