Brain-to-Language Model: Decoding Semantic Speech Representations from EEG via Cross-Modal Alignment
(2026) BMEM01 20261Division for Biomedical Engineering
- Abstract
- A key limitation of current hearing devices is the lack of user feedback to determine what speech content is actually being registered by the brain. To provide an objective measure of this cognitive processing for future brain-steered hearing devices, this study develops a brain-to-language model capable of decoding semantic-like information directly from non-invasive electroencephalography (EEG) signals. After modelling temporal response functions to confirm the correlation between Whisper embeddings and EEG signals, two distinct EEG encoders are trained via contrastive learning. Specifically, a dilated CNN architecture and a transformer model are utilised to map continuous neural data directly to the contextualised semantic embeddings... (More)
- A key limitation of current hearing devices is the lack of user feedback to determine what speech content is actually being registered by the brain. To provide an objective measure of this cognitive processing for future brain-steered hearing devices, this study develops a brain-to-language model capable of decoding semantic-like information directly from non-invasive electroencephalography (EEG) signals. After modelling temporal response functions to confirm the correlation between Whisper embeddings and EEG signals, two distinct EEG encoders are trained via contrastive learning. Specifically, a dilated CNN architecture and a transformer model are utilised to map continuous neural data directly to the contextualised semantic embeddings extracted by Whisper. Since there is no observable ground truth for human semantic processing, this study is using Whisper’s artificial embeddings to represent them. Despite this limitation, results demonstrate that meaningful semantic patterns can be extracted from EEG recordings, with both models performing above chance. The transformer architecture achieved the best results by successfully mapping neural signals to the uncompressed semantic space, outperforming the dilated CNN which required dimensionality reduction. (Less)
- Popular Abstract
- Extracting Speech and Meaning from Brain Signals
Hearing loss is a growing global challenge that demands smarter technology. This thesis investigates the use of artificial intelligence to match the patterns of brainwaves directly with the features of sound in order to decipher what a listener truly perceives and understands.
Hearing loss does not only make sounds fainter, it often distorts speech, making it incredibly exhausting for the brain to process and interpret the meaning of words. While modern hearing aids excel at amplifying sound, adjusting them still relies heavily on the user's subjective feedback and personal guesswork. This creates a critical need for smart, brain-steered hearing aids that can objectively measure how... (More) - Extracting Speech and Meaning from Brain Signals
Hearing loss is a growing global challenge that demands smarter technology. This thesis investigates the use of artificial intelligence to match the patterns of brainwaves directly with the features of sound in order to decipher what a listener truly perceives and understands.
Hearing loss does not only make sounds fainter, it often distorts speech, making it incredibly exhausting for the brain to process and interpret the meaning of words. While modern hearing aids excel at amplifying sound, adjusting them still relies heavily on the user's subjective feedback and personal guesswork. This creates a critical need for smart, brain-steered hearing aids that can objectively measure how well a person actually understands speech.
In this work, we address this problem by designing two distinct models with the purpose of extracting patterns in brain signals and comparing these with characteristics in speech that have been extracted using a known transcription model. These comparisons are then used to train the models to find stronger similarities between the brain signal and the audio signal. If a high similarity is achieved, this suggests that the brain has successfully perceived the true meaning of the speech.
However, reading the human brain is no easy task. Brainwaves recorded through electroencephalography (EEG) are notoriously noisy and heavily contaminated by muscle movement, eye blinks and external interference. Because of this large amount of background noise, the absolute accuracy of matching the brain signals to the speech data might at first glance appear low. But when compared to random guessing, our AI models performed many times better than chance. This clear margin suggests that the models are not just making lucky guesses, but have actually managed to extract language patterns from the messy brain data.
Looking ahead, this suggests that an objective measure of comprehension can be extracted from brain waves. However, an important question remains: the lack of a ground truth, meaning that we cannot be entirely sure of what language or semantic information the AI speech model actually captures. Yet, by proving that these signals can be aligned far better than chance, this work opens the door for future hearing aids to monitor the brain in real time for automatic adjustments of the hearing aid settings. (Less)
Please use this url to cite or link to this publication:
https://lup.lub.lu.se/student-papers/record/9245750
- author
- Engström, Julia LU and O'Brien, Ellie LU
- supervisor
- organization
- alternative title
- Hjärn-till-språk-modell: Avkodning av semantiska talrepresentationer från EEG via korsmodal mappning
- course
- BMEM01 20261
- year
- 2026
- type
- H2 - Master's Degree (Two Years)
- subject
- keywords
- Generative AI, Machine learning, Semantics, Contrastive learning, Hearing loss
- language
- English
- additional info
- 2026-18
- id
- 9245750
- date added to LUP
- 2026-07-02 15:02:14
- date last changed
- 2026-07-02 15:02:14
@misc{9245750,
abstract = {{A key limitation of current hearing devices is the lack of user feedback to determine what speech content is actually being registered by the brain. To provide an objective measure of this cognitive processing for future brain-steered hearing devices, this study develops a brain-to-language model capable of decoding semantic-like information directly from non-invasive electroencephalography (EEG) signals. After modelling temporal response functions to confirm the correlation between Whisper embeddings and EEG signals, two distinct EEG encoders are trained via contrastive learning. Specifically, a dilated CNN architecture and a transformer model are utilised to map continuous neural data directly to the contextualised semantic embeddings extracted by Whisper. Since there is no observable ground truth for human semantic processing, this study is using Whisper’s artificial embeddings to represent them. Despite this limitation, results demonstrate that meaningful semantic patterns can be extracted from EEG recordings, with both models performing above chance. The transformer architecture achieved the best results by successfully mapping neural signals to the uncompressed semantic space, outperforming the dilated CNN which required dimensionality reduction.}},
author = {{Engström, Julia and O'Brien, Ellie}},
language = {{eng}},
note = {{Student Paper}},
title = {{Brain-to-Language Model: Decoding Semantic Speech Representations from EEG via Cross-Modal Alignment}},
year = {{2026}},
}