Why Brain-Computer Interfaces Need Recalibration, and How Decoders Are Learning to Cope With Neural Drift
October 8, 2026 — by sysop_gray — filed under Machine Intelligence
Most coverage of brain-computer interfaces shows the best day: a cursor gliding across a screen, a sentence appearing as a paralysed person attempts to speak, a headline accuracy figure. Most people do not see the next morning. The signals the decoder learned yesterday have shifted slightly overnight. Control is a little worse, and someone, often a researcher with a laptop, has to collect fresh data and retrain the system before it works properly again.
This routine is called recalibration, and it is one of the least visible and most consequential problems in the field. A decoder that needs half an hour of supervised retraining every day can perform superbly in a research session and still fail as a medical device. A person living with paralysis needs a system that works when they wake up, without a technician.
This article explains why neural signals drift, why that breaks the machine-learning models used to decode them, and how researchers have tried to fix it. It also covers what the evidence shows so far, and why the studies cannot simply be ranked against one another.

What a decoder actually learns
A BCI decoder is a model that maps recorded neural activity to an intended output. That output might be a cursor velocity, a click, a handwritten letter or a phoneme. In intracortical systems, the input is usually the firing of neurons near each electrode. It is commonly summarised as threshold crossings, counts of the moments when the voltage on a channel dips below a set level within short time bins. In surface systems such as electrocorticography (ECoG), the input is typically the power of the signal in particular frequency bands.
The decoder can be simple or complex. Early cursor systems used a Kalman filter, a linear model that estimates movement from neural activity while smoothing over time. Speech and handwriting decoders today usually use recurrent neural networks, which process sequences and output probabilities over phonemes or characters. These are often paired with a separate language model that turns those probabilities into plausible words and sentences. These are different components doing different jobs. Calling the whole pipeline "the AI" obscures where errors come from and where drift does its damage.
Whatever its architecture, a decoder trained by supervised learning learns a fixed relationship between specific input channels and specific outputs. It learns, in effect, that when electrode 37 fires vigorously and electrode 81 goes quiet, the person is trying to move the cursor up and to the left. The trouble starts when the relationship between electrodes and intentions changes.
Why the signals do not stay put
Neural recordings from implanted electrodes are nonstationary: their statistical properties change over time. The causes operate on different timescales.
Within a session, the brain moves slightly relative to the electrodes with every heartbeat and breath, and larger shifts follow changes of posture. A neuron recorded strongly on one channel at the start of a session may be recorded weakly an hour later. Attention, fatigue and medication also change baseline firing rates.
From day to day, the set of neurons an electrode picks up changes. Some neurons drift out of range and others drift in. Even a well-placed array therefore presents the decoder with a partially different population each morning.
Over months and years, the body's response to a foreign object adds slower changes. Scar tissue forms around penetrating electrodes, electrode materials degrade, and some channels lose signal entirely. Surface and epidural electrodes avoid some of these processes but face their own, including changes in the tissue between electrode and cortex.
The user changes too. People adapt their strategies as they practise with a BCI, and the decoder's errors shape how they try to move. This co-adaptation can help, but it also means the patterns the decoder learned on day one are not the patterns the person produces on day 50.
The consequence is straightforward. A decoder trained on Monday's data is tested on Tuesday's slightly different distribution, and its performance degrades. Machine learning has a name for this, dataset shift. In most applications it is an occasional nuisance. In implanted neural interfaces it happens every day.
The cost of the standard fix
The standard response has been supervised recalibration. At the start of each session, the user performs a structured task, such as attempting to move toward cued targets or reading prompted sentences aloud in their head. Because the system knows what the user was asked to do, it has labelled data. It retrains or fine-tunes the decoder, and the session proceeds.
This works, and much of the field's best published performance relies on some form of it. It also carries costs that rarely appear in headlines. Calibration time is time not spent using the device. It usually requires a trained person to run the task, and it assumes the user has the energy to perform it, which for someone with advanced ALS cannot be taken for granted. A team at the University of California, San Francisco, reported in 2020 that daily resetting of a decoder for one participant sometimes took hours. On some days the participant could not gain control at all.
There is also a quieter cost. A device that needs daily supervised calibration cannot be easily separated from the research team that knows how to calibrate it. That is a large part of why so many impressive results remain laboratory demonstrations.
Five strategies for living with drift
Researchers have attacked the problem in several distinct ways. Each makes different assumptions, and each has been tested in different conditions.
html
<table>
<thead>
<tr>
<th>Strategy</th>
<th>Core idea</th>
<th>Where the labels come from</th>
<th>Main weakness</th>
</tr>
</thead>
<tbody>
<tr>
<td>Supervised recalibration</td>
<td>Retrain or fine-tune the decoder on fresh labelled data each session</td>
<td>Cued tasks with known targets or sentences</td>
<td>Costs the user's time and usually needs a technician</td>
</tr>
<tr>
<td>Self-recalibration from inferred intent</td>
<td>Update the decoder during normal use by inferring what the user was trying to do</td>
<td>The user's own selections, such as which key they finally chose</td>
<td>Depends on the inferences being correct; needs a task with recognisable goals</td>
</tr>
<tr>
<td>Pseudo-labelling with a language model</td>
<td>Treat the language model's best interpretation of a decoded sentence as a training label</td>
<td>A language model's correction of the decoder's own output</td>
<td>Fails if the decoder is too inaccurate to produce useful corrections</td>
</tr>
<tr>
<td>Latent-space alignment</td>
<td>Map today's neural activity onto the low-dimensional structure learned on an earlier day, so the original decoder still applies</td>
<td>None; alignment is unsupervised</td>
<td>Assumes the underlying population structure is stable; tested mainly in animals</td>
</tr>
<tr>
<td>Robust or pretrained models</td>
<td>Train on many days, perturbations or subjects so the model is already tolerant of variation</td>
<td>Large archives of previously labelled data</td>
<td>Needs large datasets that most individual users cannot supply</td>
</tr>
</tbody>
</table>Learning from the user's own choices
One of the earliest practical demonstrations of self-recalibration came from the BrainGate consortium. It was published in Science Translational Medicine in 2015 by Beata Jarosiewicz and colleagues. The idea, called retrospective target inference, was to use the logic of typing. If a participant moves the cursor and eventually selects the letter R, then the movements leading up to that selection were probably intended to reach R. Those movements can be labelled after the fact and used to update the decoder, without any separate calibration task. The system also tracked shifts in baseline neural activity during pauses and corrected for them.
With these features switched on, participants' typing performance stayed stable. With them switched off, it degraded significantly. One participant, a woman with ALS, used the system across six sessions over 42 days after initial calibration on day one, without explicit recalibration. The study was small and the task was specific, but it established that the decoder's own use can supply much of the information needed to keep it on track.
Letting a language model write the labels
In 2023, a Stanford team led by Chaofei Fan, with Francis Willett and colleagues, extended this logic to brain-to-text communication. They presented the work at the NeurIPS machine learning conference. The participant, a man with a cervical spinal cord injury in the BrainGate2 trial, had two 96-channel arrays in the hand area of motor cortex. He attempted to handwrite sentences letter by letter while a recurrent neural network decoded the intended characters.
The method, called continual online recalibration with pseudo-labels (CORP), used a language model as a stand-in teacher. After each sentence, a language model chose the most plausible version of what the decoder had output. The system treated that corrected sentence as if it were the true label and fine-tuned the decoder in the background, in about nine seconds per sentence.
The study ran for 403 days, with 15 evaluation sessions roughly a month apart. Each session compared a frozen decoder, trained once and never updated, with the continually recalibrated one. Averaged across the online sessions, the frozen decoder's word error rate was 26.51%. The recalibrated decoder's was 6.16%. Removing the language model from the recalibration loop had the largest effect of any component, which shows where the approach gets its leverage.
The authors were clear about the limits. There was one participant and one task. The method depends on the decoder being accurate enough for the language model to correct it sensibly. If the decoder's raw output is too poor, the pseudo-labels become wrong and recalibration can reinforce errors. An unsupervised alignment method they also tested offered no benefit on the handwriting task. The authors suggest that may reflect the higher complexity of the neural activity involved.
Aligning today's brain to yesterday's
A different line of work starts from an observation about neural populations. Although individual neurons drift in and out of recording, the coordinated activity of the population often occupies a relatively stable low-dimensional structure, sometimes called a neural manifold. If today's recordings can be mathematically mapped back onto yesterday's structure, the original decoder can be reused unchanged.
Alan Degenhart, Byron Yu, Aaron Batista and colleagues at Carnegie Mellon and the University of Pittsburgh published a version of this approach in Nature Biomedical Engineering in 2020. They used factor analysis, a linear method for finding low-dimensional structure. The approach kept BCI control stable through recording instabilities in monkeys and in many cases recovered control after severe disruptions. A later method, NoMAD, developed by Brianna Karpowicz, Chethan Pandarinath and colleagues, used a nonlinear model of neural dynamics for the alignment. In offline analyses of monkey recordings, it decoded isometric wrist force for about three months without noticeable degradation. In the same comparison the linear method declined faster and produced more decoding failures.
These are promising results, but they come with conditions. They come mostly from monkeys performing a single, well-practised behaviour. The methods assume that the relationship between the population's structure and behaviour stays constant, which may not hold as people learn new skills. NoMAD's alignment step is computationally demanding, which matters for implanted hardware with tight power budgets.
Fixing the decoder and letting the brain adapt
The UCSF study mentioned above took a different route. Its 2020 report in Nature Biotechnology, from Daniel Silversmith, Karunesh Ganguly and colleagues, involved one participant with tetraplegia and a 128-channel ECoG array on the surface of the brain. Rather than resetting the decoder daily, the team let it accumulate data and adapt over a long period. They then held it fixed. Performance did not decline over 44 days without retraining, and after interruptions the participant quickly regained the same control patterns.
The authors attributed this to two factors. ECoG signals, which average activity over larger areas of cortex, appear more stable from day to day than recordings of individual neurons. And with a stable decoder, the brain itself can consolidate a reliable control strategy. The study was small and covered one participant, and ECoG trades signal resolution for stability. But it showed that stopping daily resets can, in the right conditions, be part of the solution rather than the problem.

Training for variation in advance
The fifth strategy is to build tolerance into the model before deployment. In a 2016 study in Nature Communications, David Sussillo, Krishna Shenoy and colleagues trained a multiplicative recurrent neural network on months of previously recorded monkey data, deliberately adding synthetic perturbations such as simulated loss of channels. The network stayed usable under conditions that disabled a standard Kalman filter, and it became more robust as the training archive grew.
More recent research has pushed this idea toward large pretrained models of neural activity, trained across many sessions and, in some cases, many subjects. The hope is that a person's new implant could be calibrated with a small amount of their own data, as speech-recognition systems adapt to new voices. That work is active and largely preclinical, and it is too early to say how far cross-participant transfer will hold for people with different injuries, implant locations and electrode counts.
What the strongest human results actually show
The most widely reported recent BCI results come from speech decoding, and they illustrate both how far decoders have come and why their numbers must be read carefully.
html
<table>
<thead>
<tr>
<th>Study</th>
<th>Participants and interface</th>
<th>Task</th>
<th>Reported result</th>
<th>Metric and conditions</th>
<th>Relevance to drift</th>
</tr>
</thead>
<tbody>
<tr>
<td>Jarosiewicz et al., Science Translational Medicine (2015)</td>
<td>BrainGate participants with paralysis; intracortical arrays</td>
<td>Typing with a cursor</td>
<td>Performance stable with self-calibration; degraded without it</td>
<td>Online typing performance; sessions over weeks</td>
<td>Recalibration from inferred intent during normal use</td>
</tr>
<tr>
<td>Silversmith et al., Nature Biotechnology (2020)</td>
<td>1 participant with tetraplegia; 128-channel ECoG</td>
<td>Cursor control and click</td>
<td>No decline over 44 days without retraining</td>
<td>Online control; study over about six months</td>
<td>Fixed decoder with stable surface signals</td>
</tr>
<tr>
<td>Willett et al., Nature (2023)</td>
<td>1 participant with ALS; intracortical arrays</td>
<td>Attempted speech to text</td>
<td>62 words per minute; 9.1% word error rate (50 words), 23.8% (125,000 words)</td>
<td>Online decoding with a language model</td>
<td>Benchmark for speed and vocabulary, not a stability study</td>
</tr>
<tr>
<td>Fan et al., NeurIPS (2023)</td>
<td>1 participant with spinal cord injury; two 96-channel arrays</td>
<td>Attempted handwriting to text</td>
<td>Word error rate 6.16% with recalibration versus 26.51% frozen, over 403 days</td>
<td>Online averages across 15 monthly sessions</td>
<td>Pseudo-labelled recalibration without user effort</td>
</tr>
<tr>
<td>Card et al., New England Journal of Medicine (2024)</td>
<td>1 participant with ALS; four arrays, 256 electrodes</td>
<td>Attempted speech to text</td>
<td>99.6% word accuracy on day one (50 words); 97.5% over 8.4 months (125,000 words); about 32 words per minute in conversation</td>
<td>Online decoding with continued training data; over 248 hours of use</td>
<td>Rapid calibration plus accumulating data</td>
</tr>
</tbody>
</table>The 2024 New England Journal of Medicine report from UC Davis, led by Nicholas Card and David Brandman, is the clearest example of a decoder improving rather than degrading over time. A 45-year-old man with ALS received four microelectrode arrays, 256 electrodes in total, in the speech-related motor cortex. On the first day of use, 25 days after surgery and after about 30 minutes of calibration, the system decoded his attempted speech with 99.6% word accuracy on a 50-word vocabulary. After 1.4 more hours of training data on the second day, accuracy reached 90.2% on a 125,000-word vocabulary. As data accumulated, accuracy was maintained at 97.5% over the following months. By the time of the report he had used the system for more than 248 hours, conversing at about 32 words per minute.
The result is remarkable. It is also a single participant, with a decoder that kept receiving new training data, supported by a research team.
Why these numbers cannot be lined up
It is tempting to rank these studies by their headline figures. That would be a mistake, for several reasons.
The metrics are different. Word accuracy and word error rate are related, but word error rate also counts inserted words, so the two do not simply sum to 100%. Words per minute measures speed, not correctness. A cursor study's measure of stable typing performance cannot be compared with a speech study's error rate. Correlation measures (R²) in offline monkey analyses describe how well decoded signals track recorded behaviour, not whether a person could communicate.
Vocabulary size changes the task. Choosing among 50 words is a different problem from producing open text from a 125,000-word vocabulary. Error rates on the two are not comparable, even within one study.
Offline and online results differ. An offline result reanalyses recorded data and can be tuned after the fact. An online result is what the participant experienced in real time, including their own reactions to the decoder's errors. The CORP study reports both and shows they differ.
Language models do a great deal of the work. In speech and handwriting decoders, a language model corrects raw neural decoding into plausible text. In the CORP study, removing the language model from recalibration had the largest effect of any component. Reported word error rates therefore reflect the combination of neural decoder and language model, not the quality of neural decoding alone.
Almost every study involves one participant. These are proof-of-concept demonstrations in carefully selected people, with arrays in specific brain areas, supported by expert teams. They show what is possible. They do not show how often it will work across the range of people who might need a BCI.
From accuracy to independence
The deeper lesson of the recalibration problem is that accuracy and usability are different things. A decoder that reaches 97% accuracy after expert recalibration and one that reaches 90% with no help at all serve very different lives. The second may be more useful to a person at home, because it works when they need it.
That is why research on recalibration is quietly becoming central to the field's clinical future. The questions that matter for a person using a BCI are not only how accurate it is on its best day. They are also how long it takes to start working each morning, whether it needs someone else to set it up, and how it behaves when signals change without warning.
Several approaches covered here aim at exactly that. Some, like CORP and retrospective target inference, harvest labels from ordinary use. Others, like manifold alignment, stabilise the neural representation the decoder sees. Robust pretraining builds tolerance in advance, and the fixed-decoder approach lets the brain do some of the adapting. None has yet been shown in a large trial, across many participants, over years of unsupervised home use. Until that evidence exists, the most honest description of the field is that its decoders have become strikingly accurate in the laboratory, and are only beginning to become independent of it.