NEUROLINGUA (ds007111)
NEUROLINGUA: A Neuroimaging Database Tailored to Unraveling the Complexity of Multilingual Comprehension
Brain fMRI of healthy adults performing a language localizer in Spanish, Basque or English, pooled to study how the language comprehension network varies with bilingual and multilingual profiles. Includes structural MRI derivatives and sociolinguistic and cognitive measures.
Overview
NEUROLINGUA brings together fMRI language localizer runs that the Basque Center on Cognition, Brain and Language (BCBL) in Donostia-San Sebastian collected across many of its projects between 2010 and 2022. All participants lived in Gipuzkoa, where Basque and Spanish are both everyday languages, so the cohort spans people who are close to monolingual in Spanish through to balanced Basque-Spanish bilinguals and speakers of three languages. The aim is to study how the language comprehension network varies with age of acquisition, exposure and proficiency. The data descriptor appeared in Scientific Data in 2026, and the release is on OpenNeuro as ds007111.
Composition
The paper and participants.tsv list 725 healthy adults aged 18 to 82 (389 women; mean age 34, median 27). Snapshot 1.1.1 holds 1,684 BOLD runs: 717 subjects have at least one Spanish run, 294 a Basque run and 57 an English run, and many subjects have repeated runs. Derivatives include CAT12 morphometry (atlas volumes, cortical surfaces and thickness) for 723 subjects, SPM12 probabilistic atlases of the auditory and visual comprehension contrasts, and a phenotype table with sociolinguistic, education and cognitive measures. The stimuli for all three languages are included.
Acquisition
All scans come from a 3 T Siemens Magnetom scanner. T1-weighted images are 3D MPRAGE at 1 mm isotropic with a 64-channel head coil. Functional runs are gradient-echo EPI; the two most common settings are TR 850 ms with 2.4 mm voxels and TR 2400 ms with 3 mm voxels. The task is an adaptation of the Pinel localizer: auditory and visual sentences for passive comprehension, mental arithmetic or a button press, plus flashing checkerboards.
Annotations
There are no image labels. Events files give onsets and conditions of every trial.
Known limitations
- The original T1-weighted images are not on OpenNeuro. BCBL shares them only for justified requests under its own agreement, and asks researchers to use the CAT12 derivatives instead where possible.
- Data come from many projects, so acquisition parameters, number of runs and languages tested differ between people. Some subjects have more than the 14 runs the paper states as the maximum.
- Phenotype variables are missing for part of the cohort, with missing rates from 1 to 52 percent per variable.
- Three subjects have T1-weighted scans of too low quality for morphometry.
Cohort
Aggregate numbers from the sources below. Bars are relative to the 725 subjects.
Sex
- Female 389 54%
- Male 336 46%
Methods: Sample description and data harmonization in Quiñones et al. 2026, Unraveling the Complexity of Multilingual Comprehension, Scientific Data; sex column in participants.tsv of ds007111 snapshot 1.1.1 (CC0), age and sex columns
Contrast / sequence
Groups can overlap
- BOLD fMRI 725 100%
- T1-weighted 725 100%
From File tree of ds007111 snapshot 1.1.1 (CC0), NIfTI files counted per subject folder; Methods: Sample description and data harmonization in Quiñones et al. 2026, Unraveling the Complexity of Multilingual Comprehension, Scientific Data
Age
mean 34 ± 16.6, median 27, range 18 to 82
age column in participants.tsv of ds007111 snapshot 1.1.1 (CC0), age and sex columns; Methods: Sample description and data harmonization in Quiñones et al. 2026, Unraveling the Complexity of Multilingual Comprehension, Scientific Data
Age by sex
Reported cross table. Missing cells were not published (fewer than 10 or not reported).
- 1870-7926
- 3160-6918
- 1750-5920
- 2640-4928
- 7030-3964
- 18420-29152
- 4310-1926
age and sex columns in participants.tsv of ds007111 snapshot 1.1.1 (CC0), age and sex columns
Contrast combinations
How many subjects have exactly each set of contrasts.
| T1w | bold | Subjects with exactly this set |
|---|---|---|
725 |
Methods: Sample description and data harmonization in Quiñones et al. 2026, Unraveling the Complexity of Multilingual Comprehension, Scientific Data
License and access
Our reading of the license, not legal advice. Before you use the data, read the original license and confirm that your use is allowed. We take no responsibility for how you use a dataset. Full disclaimer
Download without an account
The BOLD runs, events files, stimuli, CAT12 structural derivatives, SPM12 group atlases and phenotype tables are an open download on OpenNeuro. The original T1-weighted images are not in the OpenNeuro release; BCBL shares them only for justified requests through a web form, after acceptance of a BCBL data use agreement.
Different parts of the data carry different licenses. The summary on the right shows the most restrictive answer per rule.
Creative Commons Zero 1.0 Universal
Public domain dedication. Do anything with the data, including commercial use, without asking and without having to give credit.
dataset_description.json states "License" CC0. The README usage notes add that the data are for academic researchers, institutions and entities on condition of proper attribution, and that the authors' names may not be used to endorse commercial products without written permission.
What you can do
- Yes
- Yes
- Yes
- Yes
- Yes
What you can share
- Yes
- Yes
- Yes
What you must do
- No
- Share alike No
- No
- No
- Manuscript review No
- Release code No
- Return results No
- Delete after use No
Limits
- No
- Location limits No
No published license or data use terms
The data holder has published no license and no data use terms. Any use, sharing or commercial right has to be agreed with the data holder, and nothing can be assumed to be allowed. Default copyright and data protection law still apply.
The BCBL page requires a scientific justification and acceptance of a BCBL data use agreement, whose text is not published.
What you can do
- Commercial use Not stated
- Not stated
- Train ML models Not stated
- Create derived data Not stated
- Publish results Not stated
What you can share
- Share the data Not stated
- Share derived data Not stated
- Share trained models Not stated
What you must do
- Cite or credit Not stated
- Share alike Not stated
- Sign an agreement Not stated
- Ethics approval Not stated
- Manuscript review Not stated
- Release code Not stated
- Return results Not stated
- Delete after use Not stated
Limits
- No re-identification Not stated
- Location limits Not stated
Citation
Quiñones I, Carrión-Castillo A, Diez-Zabala I, et al. Unraveling the Complexity of Multilingual Comprehension: Neuroimaging and Linguistic Profiling in 700+ Adults. Scientific Data 13, 1315 (2026). doi:10.1038/s41597-026-07423-9
Sources
Every number on this page comes from one of these documents. Each chart names the table or page it is taken from. The raw numbers are in stats.csv.
- Quiñones et al. 2026, Unraveling the Complexity of Multilingual Comprehension, Scientific Data paper
- OpenNeuro ds007111 snapshot 1.1.1 (dataset_description.json, README and snapshot summary) website
- participants.tsv of ds007111 snapshot 1.1.1 (CC0), age and sex columns computed
- File tree of ds007111 snapshot 1.1.1 (CC0), NIfTI files counted per subject folder computed
- BCBL Data Sharing page for NEUROLINGUA website