Notes
The Speech from first principles posts are long and linear. This is the other way in: one concept per page, each standing on its own and linking to every post and note that leans on it.
Every page ends with a Referenced by list, built automatically from what actually links to it.
-
All-pole filterLPC · linear predictive coding
The transfer function of a lossless tube chain: poles, no zeros. Why LPC works — and where it breaks.
-
CTCConnectionist Temporal Classification
Connectionist Temporal Classification — training a recogniser to align frames to labels by summing over every possible alignment.
-
Cepstrum
The inverse transform of the log spectrum — separating the slow formant envelope from the fast pitch comb by a second Fourier step.
-
Formantresonance · F1 / F2 / F3
A resonance of the vocal tract — a peak in the spectral envelope that carries vowel identity, independent of pitch.
-
Glottal sourceglottal flow · the source
The buzz from the vibrating vocal folds — a harmonic comb carrying pitch and voice quality, but no vowel.
-
Mel filterbankMFCC front end · log-mel
A bank of triangular filters on a perceptual frequency scale — the cheap, reliable way to get the spectral envelope for recognition.
-
Reflection coefficientPARCOR · k
The single number governing a step in tube area. Chain them into a vocal tract; fit them to a signal and they are PARCOR coefficients.
-
Source–filter model
Speech = a source spectrum × a vocal-tract filter × lip radiation. The factorisation the whole field rests on.
-
Vocal tracttract
The ~17 cm acoustic tube from glottis to lips whose shape sets the formants. Modelled as a chain of cylinders.