Speech neuroprosthesis is a brain–computer interface that recovers intended speech from neural activity and renders it as text, synthesized voice, or an animated face. It is aimed at people who retain the intention and the neural machinery for speech but have lost the motor pathway: those with amyotrophic lateral sclerosis, brainstem stroke, or severe anarthria from other causes. Between 2021 and 2025 the field moved from decoding a fifty-word vocabulary at a conversational crawl to near-real-time output over vocabularies of a hundred thousand words.
Crucially, these systems decode attempted articulation, not thought. The signals come from the region of cortex that would drive the lips, jaw, tongue, and larynx; the decoder recovers a motor plan, and language models convert that noisy plan into plausible sentences. That distinction governs both what the technology can do and what privacy risks it does and does not create.
How it works
Speech production has a somatotopic map. Along the ventral part of the precentral gyrus, cortical populations encode the movements of individual articulators, and this organization persists in people who have been unable to speak for years — the map does not disappear when the output pathway fails. That persistence is the enabling fact for the whole field.
A speech neuroprosthesis records from that cortex, extracts features (high-gamma power for Electrocorticography interfaces, threshold crossings and band power for a penetrating Utah array), and passes them to a sequence model that emits phonemes or subword units. A language model then converts the phoneme stream into words, using the statistics of English to resolve ambiguity that the neural signal alone leaves open. These are among the few clinical research systems in which a statistical language model of the kind central to debates about Artificial general intelligence sits inside a control loop with a person's nervous system. The user attempts to speak, silently or with whatever residual movement remains, and the pipeline runs in near real time.
Output can be text on a screen, a synthesized voice, or both. Some systems personalize the voice from recordings made before the person lost speech, and one has driven a digital avatar whose face moves with the decoded articulation. This is the same Neural decoding architecture used for cursor control, applied to a much higher-dimensional output space.
The language model does a lot of the workRaw phoneme decoding accuracy is well below the headline word accuracy. A substantial share of performance comes from the language model choosing the most probable sentence consistent with a noisy neural signal. This is legitimate engineering — human listeners do the same thing — but it means the reported word error rate is a property of the whole system, and that unusual or unpredictable utterances are decoded worse than the aggregate figure suggests.
Development history
The 2021 report that established feasibility involved a man with anarthria following a brainstem stroke, implanted with a surface grid over sensorimotor cortex. He produced sentences from a fifty-word vocabulary at roughly fifteen words per minute, with about a quarter of words wrong.1 Slow, error-prone, and unmistakably speech.
Two 2023 papers changed the scale. A Stanford group using four penetrating arrays in a participant with ALS reported 62 words per minute over a 125,000-word vocabulary, with a word error rate near 24 percent — and under 10 percent when the vocabulary was restricted to fifty words.2 A UCSF and Berkeley team, using a 253-channel surface array in a woman paralysed by a brainstem stroke, reported a median 78 words per minute over a 1,024-word vocabulary, together with a synthesized voice built from pre-injury recordings and an avatar that moved its face.3
A 2024 report from a UC Davis and Brown collaboration pushed accuracy rather than speed, reporting word accuracy near 98 percent over a very large vocabulary in a participant with ALS, with the system usable from the first session and improving over months of daily self-directed use.4 Subsequent work has attacked latency: streaming architectures that emit words as they are formed rather than at the end of a sentence, and voice synthesis fast enough to preserve conversational turn-taking and some prosody.5
A related line uses imagined handwriting rather than speech. Decoding attempted pen strokes from motor cortex produced roughly 90 characters per minute in one participant, which remains the fastest character-level result and demonstrates that the choice of imagined movement is itself a design parameter.6
What has and has not been shown
Shown: fluent, large-vocabulary decoding in individual participants; stable enough performance for daily use over months; voice synthesis with recognisable identity; decoding from both surface and penetrating electrodes; usable results from the first day of calibration in at least one case.
Not shown: any of this in more than a handful of people. Nearly every result above comes from a single participant, and different participants have different implant sites, different disease stages, and different residual capabilities. There is no multi-participant efficacy trial, no approved device, no data on how the systems perform in noisy real-world settings with fatigue and changing medication, and almost no work outside English. Most systems remain tethered through a percutaneous connector to laboratory hardware.
A further limitation is that these decoders address speech production. They do not help people whose language comprehension or formulation is impaired, which excludes many stroke and dementia patients. A Cochlear implant restores an input channel; a speech neuroprosthesis restores an output channel; neither touches the linguistic system between them.
Within the wider field of Neuroprosthetics, speech decoding is unusual in producing an output that is unambiguously interpretable. Efforts at cognitive prosthetics further upstream — the hippocampal work described under Memory prosthesis — report effects that are small, contested, and difficult to evaluate precisely because there is no equivalent of a transcript to score.
Inner speech and privacy
The obvious worry — that a device might transcribe thoughts a person did not choose to utter — is neither science fiction nor imminent. Attempted speech and imagined speech produce overlapping but distinguishable patterns in motor cortex, and a 2025 study reported that imagined speech is decodable from the same electrodes, while also demonstrating a keyword-gated mode designed so that the decoder ignores neural activity until the user internally produces a chosen password.7
That is a reassuring engineering answer to a real problem, and it establishes the shape of the issue: the safeguard has to be designed in, because the signal is there whether or not the system is asked to use it. The broader questions belong to Mental privacy and the developing Neurorights frameworks — who holds the neural recordings, whether decoded text has legal status as speech, and what happens when a decoder outputs a sentence the user did not intend. Error correction in a speech prosthesis is not a cosmetic matter; a wrong word attributed to a person who cannot verbally repudiate it is a distinctive harm.
Outlook
The near-term path is a pivotal trial in ALS, the population with the clearest need and the strongest existing results, using a fully implanted wireless system rather than a percutaneous one. Progress there depends less on decoding accuracy, which is already adequate, than on hardware longevity, surgical throughput, and the unglamorous work of making a device that a clinician can implant and a family can support at home — the same constraints that shape every implanted Brain–computer interface, from the endovascular Stentrode to the penetrating threads of Neuralink.
These systems are also the strongest existing evidence for the modest version of Human–AI merger: a machine-learning model sitting between cortex and the world, doing inferential work the user's damaged nervous system cannot, with the user unable to distinguish their own contribution from the model's.
The scientific question underneath is how much of language is reachable this way at all. Decoding articulation works because articulation is motor. Semantic content, the words a person has not yet converted into a motor plan, is represented in distributed cortical patterns that no implanted electrode array currently samples, and it is not known whether a device could recover it without covering far more of the brain than any surgery would justify.
See also
- Brain–computer interface
- Electrocorticography interfaces
- Utah array
- Neural decoding
- Neuroprosthetics
- Mental privacy
- Neurorights
- Brain-to-brain interfaces
References
Footnotes
-
paperMoses, D. A. et al. "Neuroprosthesis for decoding speech in a paralyzed person with anarthria." New England Journal of Medicine, 2021. ↩
-
paperWillett, F. R. et al. "A high-performance speech neuroprosthesis." Nature, 2023. ↩
-
paperMetzger, S. L. et al. "A high-performance neuroprosthesis for speech decoding and avatar control." Nature, 2023. ↩
-
paperCard, N. S. et al. "An accurate and rapidly calibrating speech neuroprosthesis." New England Journal of Medicine, 2024. ↩
-
paperLittlejohn, K. T. et al. "A streaming brain-to-voice neuroprosthesis to restore naturalistic communication." Nature Neuroscience, 2025. ↩
-
paperWillett, F. R., Avansino, D. T., Hochberg, L. R., Henderson, J. M. and Shenoy, K. V. "High-performance brain-to-text communication via handwriting." Nature, 2021.↩The participant had a spinal cord injury rather than a speech disorder, so the result tests decoding rather than the clinical use case for a speech device.
-
paperKunz, E. M. et al. "Inner speech in motor cortex and implications for speech neuroprostheses." Cell, 2025.↩The password-gated mode is a laboratory demonstration in already implanted research participants, not a safeguard validated in any deployed device.