7-9 rue de l’Atlas - Paris, France

Projects

Open-sourcing the modeling of brain activity for science and medicine

The Digital Brain Project aims to develop a neural encoding model. Over the past several years, one of the laboratory’s primary research directions has been neural decoding, namely the prediction of mental representations from brain recordings, whether invasive (using intracranial electrodes) or non-invasive (using functional MRI or magnetoencephalography). By reversing the inputs and outputs of such models, it becomes possible to perform the complementary operation: neural encoding. In practice, this involves predicting how a human brain should respond to a given stimulus. In other words, the goal is to build digital twins of the brain. Highly encouraging results have recently been obtained using the TRIBE v2 model proposed by Stéphane d’Ascoli (Brain AI team) for functional MRI data. However, the most suitable data for training such models differ substantially from the traditional datasets commonly available in cognitive neuroscience. Indeed to train suc a model, collecting data from a large number of participants performing a cognitive task for one or two hours is considerably less informative than acquiring several dozen—or even one hundred—hours of data from a smaller number of individuals.

With this objective in mind, the Digital Brain Project seeks to launch a large-scale international initiative that would actively fund a limited number of world-class research teams tasked with recording a relatively small number of brains over very long periods of time. Through close collaboration, new model architectures will be developed and evaluated in order to achieve increasingly accurate neural encoding. The potential impact of this effort is considerable. Such models would make it possible to conduct simulated cognitive science experiments and reserve costly human recordings for the most promising findings. They could also enable major advances in brain–computer interfaces by providing a precise understanding of the information that should be transmitted to the brain trough invasive electrodes. Furthermore, models trained on healthy brains could be adapted to address pathological conditions, paving the way toward patient-specific precision medicine in numerous domains, including epilepsy, inflammatory neurological disorders, stroke, and other neurological conditions.

How the Brain Learns Language: From Sounds to Words

How does a child's brain come to process language?

To find out, we recorded brain activity from 46 French-speaking participants (aged 2 to 46) undergoing clinical treatment for epilepsy as they listened to Le Petit Prince. Using over 7,400 intracranial electrodes, we tracked how the brain represents phonetic and lexical features across developmental ages.

Our results reveal that children aged 2–5 already show clear language representations in the auditory cortex. However, while phonetic processing remains stable over time, lexical representations strengthen significantly with age, progressively expanding into higher-order areas like the frontal and anterior temporal cortices.

Finally, comparing this trajectory to AI, we found that training language models (wav2vec 2.0 and Llama 3.1) spontaneously aligns them with human brain data. Notably, Llama 3.1’s learned features mirror the advanced language maturation seen in older participants but absent in toddlers, highlighting AI as a powerful tool for modeling human cognitive development.

Language decoding using micro-electrophysiological data

Deciphering Language from Single Neurons to Large Networks

At the hospital, we explore the frontiers of human language by partnering with patients undergoing clinical monitoring for severe epilepsy. To safely locate the source of their seizures, these patients have specialized stereo-electroencephalography (SEEG) electrodes temporarily implanted deep within their brains. For our research, we utilize advanced “micro-macro” versions of these sensors. While the “macro” contacts track how large language networks communicate across the brain, hair-thin “micro” wires allow us to listen to the high-resolution electrical chatter of individual neurons.

As patients volunteer to participate in our studies while reading, listening, or speaking, we record this rich, multi-layered electrophysiological data to advance “language decoding”—translating microscopic electrical sparks into specific sounds, syllables, and words. To crack this complex code, we feed these massive datasets into sophisticated machine learning algorithms. Ultimately, our goal is to train AI on this unprecedented level of detail to build a micro-scale foundational model for brain cognition, revealing exactly how the building blocks of human thought and language are physically orchestrated within the mind.

Making Better Use of Brain Data: From Minutes to Days: Scaling Intracranial Speech Decoding with Supervised Pretraining

From Minutes to Days: Scaling Intracranial Speech Decoding with Naturalistic Pretraining

Traditional speech decoding research typically relies on participants sitting still through brief, rigidly controlled experiments. While this provides clean data, the sheer volume is highly limited. In this study, we took a radically different approach: leveraging the vast amount of brain activity and ambient audio recorded continuously, 24 hours a day, during patients’ week-long hospital stays for epilepsy monitoring.

By utilizing roughly 100 times more data than standard experiments, we used these continuous recordings of real-world interactions, television, and background noise to pre-train a computational model. When tested against a controlled audiobook task, the results were clear: models pre-trained on week-long naturalistic data significantly outperformed those trained only on standard experiments. Furthermore, performance scaled consistently with data volume, showing no signs of plateauing.

While a brief period of fine-tuning remains necessary to bridge the gap between noisy hospital environments and clean experimental audio, this approach unlocks an unprecedented resource. Despite challenges like day-to-day drift in neural responses, this work proves that clinical continuous monitoring data can powerfully accelerate our understanding of speech processing and the development of robust brain-computer interfaces for communication restoration.

Understanding the Mind as It Types: From Neural Dynamics to Brain-Computer Interfaces

From Thought to Action: Mapping Hierarchical Language Production for Non-Invasive BCIs

How does a thought become a sentence? Using non-invasive magnetoencephalography (MEG), our lab tracked healthy volunteers typing sentences to map the neural cascade from abstract meaning down to precise motor execution.

We discovered that language production is strictly hierarchical: the brain processes multiple linguistic levels simultaneously (context, words, syllables, letters), with broad semantic meaning activating long before individual keystrokes are executed.

Leveraging this mechanistic insight, we built a deep learning pipeline that successfully decodes typed characters directly from non-invasive MEG signals. This marks the first time character-level brain-to-text decoding has been achieved without surgery, paving the way for safe, accessible communication-restoration technologies.