Closing the gap between thought and generative speech with Deep Learning

Vocal communication skills develop in infancy and are used daily to convey emotion, share ideas, and express needs. Instantaneous exchanges flow effortlessly from one person to another. However, loss of speech due to neurological injury limits an individual’s ability to easily take part in this integral aspect of human connection. Current assistive communication devices (e.g. text-to-speech interfaces, speech synthesisers) cause speech delays and interruptions that disrupt the natural cadence of dialogue. This loss of normal speaking rhythm diminishes the quality of spoken communication, increasing feelings of frustration and isolation. To address gaps between thought and speech production, Littlejohn et al designed and evaluated a novel brain-to-speech neuroprosthesis.
Using Deep Learning to generate speech from brain activity
The specific region of the human brain that allows vocalisation of thoughts is the speech sensorimotor cortex. Although previous studies have demonstrated that information encoded in this region can generate text or synthesised speech, researchers relied on explicit vocalisations of healthy participants to ‘train’ the speech synthesisers. Therefore, applicability to individuals with impaired vocal ability was uncertain. The synthesised speech was not produced until complete sentences were generated, causing speech delays and lack of continuity in the final audible output. Consequently, the delay in generated speech increased with the length or complexity of the spoken attempt.
In contrast, Littlejohn et al used deep learning to develop a neuroprosthesis that:
- Did not rely on any form of vocalisation for training, and
- Could produce continuous, spontaneous speech without the need for external prompts.
To achieve these goals, the neural activity of a study participant who had lost vocal ability after a brainstem stroke was monitored as they attempted to mouth sentences containing words from pre-defined vocabularies. The participant was implanted with an array of 253 electrodes over their speech sensorimotor cortex that measured changes in electrical impulses by electrocorticography (i.e. directly from the surface of the brain). The system then used deep learning to decode the signals in short (80 ms) windows. The rapid decoding process minimised delays in speech synthesis, allowing audible speech and text outputs to be generated continuously rather than waiting until the participant completed their speech attempt. Pre-injury recordings of their voice were used to personalise the generated speech, incorporating a natural speaking rhythm.
Study findings
The brain-to-speech neuroprosthesis was evaluated by measuring its ability to decode phrases and sentences from a small-vocabulary ‘50-phrase-AAC’ set and a large vocabulary ‘1,024-word-General’ set, the latter of which contained 12,379 unique sentences. The framework generated both audible speech and text decoding in less than 2 seconds from the onset of the speech attempt. The rate of word decoding was nearly twice as fast as previously reported systems. For example, a median of 47.5 words per minute (WPM) were decoded from the 1,024-word-General set, compared with 28.3 WPM using the previous approach. Speech and text generation was particularly successful with the 50-phrase set, which included vital primary caregiving phrases. Error rates were under 15% for both generated speech and decoded text for this set of phrases. The model was also successful in decoding novel sentences (using words from the 1,024-word-General set) and novel words (outside of the training sets) at above chance probabilities.
The researchers then applied their decoding framework to different silent-speech interfaces. These included the additional open-vocabulary datasets from: 1) a paralysed participant with a brain-computer interface (microelectrode array) implanted on the surface of their brain and 2) a healthy participant with non-invasive surface electrodes placed over their vocal tract (electromyography). Each modality achieved above-chance performance on silent speech attempts. Overall, this analysis showed that the decoding framework could be applied to other silent-speech interfaces, but more trials are needed to ensure that speech delays and errors are minimised.
Not so fast
Despite the significant improvement in response time, reduced latency came at the cost of accuracy. Error rates for sentences derived from the 1024-word set were much larger compared to the 50-phrase set, with word-error rates approaching 60%. This could be attributed to the larger vocabulary of the 1024-word set compared to the 50-phrase set, which only contained 119 unique words. There may also have been bias towards the 50-phrase set during training. Text decoding generally performed better across all trials, highlighting a limitation in the speech synthesis itself. The model’s ability to decode words and sentences outside of the training sets also showed lower accuracy and results were less consistent. Finally, as the device was trained and assessed on only a single participant, its complete generalisation to a larger population was not evaluated.
Why does it matter?
The neuroprosthesis developed by Littlejohn et al produced intelligible, audible speech and text more rapidly than current state-of-the-art speech synthesis devices. This task was accomplished without the need for participant vocalisation at any stage of training or evaluation. Overcoming dependence on vocalisations by decoding neural signals alone and reducing latencies in speech generation would restore natural communication to individuals with vocal tract paralysis. The user has volitional control over their communication, and the personalisation of speech output to match their pre-injury voice and speaking rhythm helps them to regain a sense of identity. The work shows promising progress towards the goal of restoring accurate, fluent speech for patients with severe paralysis.
Take home messages
- A brain-computer interface that uses small windows of neural activity and deep learning to generate speech has been developed.
- The device allowed near-instantaneous (<2s), volitional speech synthesis for a user with vocal tract paralysis.
- With minor adjustments the models can be successfully applied to other silent-speech devices, including non-invasive technologies.
Guest author: Eleni Nestoros, PhD
Reviewer: Barbara Fahmy, MS OTR, MPA
This article was written as part of a series of ‘journal club’ summaries for Scientific Writers Ltd and is based on the following publication.
Title: A streaming brain-to-voice neuroprosthesis to restore naturalistic communication
First Author: Littlejohn, KT
Journal: Nature Neuroscience
Date online: 31 March 2025
Other references
None
TL;DR
Scientists develop a brain-to-speech device that combines brain activity with deep learning models to generate audible speech and text in patients with severe paralysis.





