New audio tokenization method improves how AI models process human speech

New audio tokenization method improves how AI models process human speech

Researchers have developed a way to compress complex speech data into smaller units, allowing multimodal systems to understand sound without losing emotional nuance or quality.
GP
Giulio Prisco
Dec 4, 2025
2 min read

Large language models (LLMs) were originally created to process and generate text, and then evolved into multimodal systems capable of understanding and producing multiple types of media at the same time, such as images, music, and speech. While text is relatively easy for computers to handle, integrating spoken language remains a significant technical hurdle.

Currently, the standard method for adding speech to these models involves converting sound waves into audio tokens. These tokens are small digital building blocks that function for sound the way the alphabet functions for writing. However, human speech is incredibly complex. Beyond just the words spoken, a voice carries information about the speaker’s emotions, their unique identity, and their accent. Because of this richness, standard audio tokens usually have a high bitrate, which refers to the large amount of digital information contained in each second of recording. This data density makes it very difficult for artificial intelligence (AI) to learn from speech quickly or efficiently.

Shrinking data without losing quality

To solve this problem, researchers developed a new tool called FocalCodec. This technology creates audio tokens that are much more efficient than previous versions. It achieves this by using a method called binary spherical quantization, which converts complex audio signals into compact, simple units. Additionally, it utilizes a technique known as focal modulation to help the computer system concentrate on the most important parts of the speech data.

The group verified their results by asking 33 participants to listen to various audio samples. These listeners compared original recordings against speech that had been compressed and reconstructed by the new software. The participants found the results to be nearly identical to the originals, proving that the system can significantly reduce data size without making the audio sound distorted or robotic. The authors believe this technology will eventually allow AI to understand sound with the same high level of proficiency it currently applies to written text.

This work has been accepted for presentation at the 39th Annual Conference on Neural Information Processing Systems. FocalCodec is described on Github.

About the Writer

More from Mindplex

Keep reading

Three more ideas worth your time.

Browse News

Discussion

Join the discussion

Sign in to share a response with the community.

Type @ to mention someone Type / or use + to add a block Highlight text, then choose Link
Loading editor

Comments cannot be edited after posting because they become part of the reputation record. Give yours a quick review first.