Notes

The Speech from first principles posts are long and linear. This is the other way in: one concept per page, each standing on its own and linking to every post and note that leans on it.

Every page ends with a Referenced by list, built automatically from what actually links to it.

  • All-pole filterLPC · linear predictive coding

    The transfer function of a lossless tube chain: poles, no zeros. Why LPC works — and where it breaks.

  • CTCConnectionist Temporal Classification

    Connectionist Temporal Classification — training a recogniser to align frames to labels by summing over every possible alignment.

  • Cepstrum

    The inverse transform of the log spectrum — separating the slow formant envelope from the fast pitch comb by a second Fourier step.

  • Formantresonance · F1 / F2 / F3

    A resonance of the vocal tract — a peak in the spectral envelope that carries vowel identity, independent of pitch.

  • Glottal sourceglottal flow · the source

    The buzz from the vibrating vocal folds — a harmonic comb carrying pitch and voice quality, but no vowel.

  • Mel filterbankMFCC front end · log-mel

    A bank of triangular filters on a perceptual frequency scale — the cheap, reliable way to get the spectral envelope for recognition.

  • Reflection coefficientPARCOR · k

    The single number governing a step in tube area. Chain them into a vocal tract; fit them to a signal and they are PARCOR coefficients.

  • Source–filter model

    Speech = a source spectrum × a vocal-tract filter × lip radiation. The factorisation the whole field rests on.

  • Vocal tracttract

    The ~17 cm acoustic tube from glottis to lips whose shape sets the formants. Modelled as a chain of cylinders.