MENU
Vowel Timbre in Spatial Audio Effect Design

Vowel Timbre in Spatial Audio Effect Design

 

Immersive audio has evolved rapidly in recent years, yet one aspect of spatial processing has remained surprisingly unexplored: the timbral character of multi channel reverberation. Most spatial reverbs and decorrelation tools are still designed to be transparent, efficient and statistically independent. These qualities are technically valuable, but they also limit the creative potential of reverb as an expressive element. In our recent peer reviewed study, we explored a different direction that brings together speech science and spatial audio design. To our knowledge, no existing spatial reverb system uses vowel formants as a timbral basis for decorrelation, and this gap opened the door to a new type of reverb architecture. By shaping multi channel reverberation with LPC derived vowel filters and velvet noise decorrelation, we show how vocal resonances can become an expressive and controllable layer within immersive sound design.

This article gives an accessible overview of the idea, why it matters and how it can be used in real world production and research.

 

Why vowels are interesting for reverb design

Every vowel has a unique spectral shape created by the resonances of the vocal tract. These resonances, often described through formants, are what make one vowel sound different from another. In speech synthesis, these shapes are commonly modeled using Linear Predictive Coding, which captures the spectral envelope of a vowel in a compact and efficient way.

In this research, these vowel derived spectral shapes are used not for speech synthesis but for shaping the timbre of multi channel reverb tails. The idea was partly inspired by choral acoustics. In large spaces, overlapping vowel sounds blend into evolving textures that become part of the reverberant field. This natural behaviour suggested a creative opportunity for immersive audio design.

 

The core method in simple terms

The system combines three main elements.

  1. A mono Feedback Delay Network reverb. This provides the initial mono reverb tail. 
  2. LPC based vowel filters. Short recordings of vowel sounds are analysed to extract their spectral envelopes. These envelopes become filters that imprint vowel like timbral characteristics onto the reverb tail.
  3. Velvet noise based decorrelation. Velvet noise is a sparse and efficient signal used in modern decorrelation and reverb design. By filtering velvet noise through vowel derived LPC filters and distributing the results across channels, the system produces a 5.1 surround field with distinct timbral identities in each channel.

Each channel receives a blend of two vowel inspired textures. This creates a spatial field that is diffuse, enveloping and subtly coloured, without becoming intelligible or distracting.

5.1 Vowel Reverb

Fig 1. A simplified overview of the vowel based reverb system used in the research. 

 

Creative possibilities

This approach allows sound designers and engineers to:

  • Transform mono into immersive space - Advanced Mono-to-5.1 surround upmixing engine with user-uploaded timbral customization
  • Create personalised 5.1 surround reverb effects - Users can record their own vowel sounds and use them to shape the reverb. Every project can have a unique spatial signature.
  • Blend timbre and space in new ways - Reverb becomes a controllable and expressive layer rather than a neutral background.
  • Explore voice-inspired textures - The method can produce reverb tails that feel organic, evolving and subtly human. This is useful in music, film, VR and installations.
  • Maintain technical robustness - Despite its creative flexibility, the system preserves effective decorrelation and wide spatial diffusion

These possibilities become even clearer when imagined in context. Because the system is shaped by user-recorded vowels, the reverb can adapt to the creative intent of each project in ways that traditional algorithms cannot. A VR forest, for example, could use a soft, user-defined /u/ shaped reverb to make open spaces feel warmer and more enveloping, with the 5.1 decorrelated field adding a sense of natural diffusion. A film scene set in a cathedral might blend the director’s own recordings of /a/ and /e/ to create a reverb tail that subtly echoes the human voice without ever becoming intelligible. In interactive installations, artists can assign different vowel-based textures to different zones, giving each area its own spectral identity and allowing the space to guide the listener through sound alone.

 

Why this matters for industry and research

This work brings together DSP development, creative sound design and immersive production practice in a way that is relevant to several groups.

  • Audio companies: a novel algorithmic concept that could inspire new plugin features or spatial processing tools.
  • Studios and sound designers: a fresh approach to shaping reverb that goes beyond conventional parameters.
  • University labs and researchers: a framework that combines speech modeling, spatial audio and perceptual evaluation. It is suitable for further exploration or collaboration.

Further Reading and Full Research Article


If you’re interested in the complete technical foundation behind this work, the full peer-reviewed study is available:

Pizzi, M., & Mróz, B. (2026). Using Vowel Characteristics for Multi-channel Signal Decorrelation and Reverberation. International Journal of Electronics and Telecommunications, 1–9.

Full text available here: https://ijet.pl/index.php/ijet/article/view/10.24425-ijet.2026.157864

Learn more about Dr Mróz: https://bmroz.eu/

The article presents the full architecture of the system, detailed implementation notes, listening test methodology, statistical analysis and references to related research. It also discusses the broader implications of using vowel-derived timbre in spatial audio, including how user-recorded vowels can be transformed into personalised 5.1 reverberation effects.
For readers who want to dive deeper into the DSP, the perceptual modelling or the creative design aspects, the publication offers a complete walkthrough of the ideas introduced in this blog. It’s a great starting point for researchers, developers and sound designers interested in exploring vowel-based spatial processing or extending the method into new formats and applications.