Vocoder · One Sound Controls Another
A vocoder sounds complicated until we open it. One sound is measured frequency-band by frequency-band. Those measurements become control signals. Then they multiply matching bands of another sound.
1 · You already know every part
CARRIER → MATCHING FILTER → × ENVELOPE → SUM → OUTPUT
Now copy that little machine many times in parallel.
2 · Play the vocoder
MODULATOR · the spectral shape
No file loaded · built-in demo ready
CARRIER · the sound being shaped
Keyboard · hold notes or a chord
Computer keys: A W S E D F T G Y H U J K · or click/touch the keyboard.
3 · Watch the voice become control information
Each meter is the envelope measured from one modulator band. The same envelope controls the matching carrier band.
Try 4 bands, then 8, 16, 32 and 64. More bands describe the changing spectral envelope in finer detail. It is the same idea repeated at progressively finer frequency resolution.
4 · Why attack and release matter
Fast envelope
Tracks consonants and rapid changes more closely. The carrier articulation becomes sharper and speech tends to become easier to recognise.
Slow envelope
Smears the measurements through time. The carrier follows a smoother spectral shape and the result becomes more pad-like.
That's the same envelope follower we already used for dynamics and modulation. It has simply been duplicated across frequency.
5 · One band is simple · a vocoder is lots of them
| Stage | What happens | Old idea |
|---|---|---|
| Split modulator | Voice enters many band-pass filters | FILTER |
| Measure | Energy in each band becomes an envelope | MEASURE / ENVELOPE |
| Split carrier | Synth enters matching band-pass filters | FILTER |
| Apply control | Each carrier band is multiplied by its voice envelope | × MULTIPLY |
| Recombine | All controlled carrier bands are added together | Σ ADD |
FILTER → MEASURE → ENVELOPE → ×
FILTER → MEASURE → ENVELOPE → ×
⋮
Σ → OUTPUT
6 · From filter bands to FFT bins
A traditional vocoder uses a bank of filters. But we have just spent several pages learning another way to represent frequency in many discrete slots.
The implementation changes, but the underlying thought is familiar: measure spectral information from one sound and use it to transform another.
What if the spectral data does not come from sound at all?
Draw a picture. Hear it. Export it. Then open the audio in a spectrogram and see the picture come back.