The artificial intelligence (AI) revolution transforms lives using large language models like ChatGPT. These models predict words based on statistical patterns in language. However, they miss non-verbal communication, such as the melody of speech.
A new study from the Weizmann Institute of Science, published in Proceedings of the National Academy of Sciences, shows that this melody, called prosody, acts like a separate language in English conversations.
Prosody includes changes in pitch, loudness, tempo, and sound quality. Pitch is how high or low a voice sounds. Loudness shows emphasis. Tempo is the speed of speech. Sound quality includes things like whispering. Prosody adds meaning beyond words, showing attitudes like curiosity or surprise. For example, a pause can change a sentence’s meaning, like “Let’s eat, Grandma” versus “Let’s eat Grandma.” Even chimpanzees and whales use prosody in their communication. In humans, it shapes whether a sentence is a question or a statement.
A dictionary of speech melodies
Researchers studied prosody by analyzing audio recordings of phone and face-to-face conversations. They used AI to identify about 200 common melodic patterns, or prosodic “words.” Each pattern, lasting about a second, carries specific meanings. For instance, a sharp rise and quick drop in pitch shows enthusiasm or agreement. The study also found rules for how these patterns combine, like simple sentences. One pattern follows another based on short-term memory, forming pairs that express single ideas, such as giving positive feedback.
Unlike spontaneous speech, scripted speech, like audiobooks, uses longer patterns and lacks these simple pairs. Prosody varies by age, social status, or historical events. It also plays a role in internal thoughts and robotic voices. The researchers aim to create an automated prosody dictionary for all languages. Future AI could use this to understand emotions or attitudes from speech melodies, improving devices like Siri or brain implants that convert thoughts to speech. This study highlights prosody’s role in human expression, paving the way for AI that captures the full depth of communication.
In related news, researchers found that listeners use gestures to predict upcoming words.