Skip to content
MRI Brain

NEUROLINGUA (ds007111)

NEUROLINGUA: A Neuroimaging Database Tailored to Unraveling the Complexity of Multilingual Comprehension

Brain fMRI of healthy adults performing a language localizer in Spanish, Basque or English, pooled to study how the language comprehension network varies with bilingual and multilingual profiles. Includes structural MRI derivatives and sociolinguistic and cognitive measures.

Overview

NEUROLINGUA brings together fMRI language localizer runs that the Basque Center on Cognition, Brain and Language (BCBL) in Donostia-San Sebastian collected across many of its projects between 2010 and 2022. All participants lived in Gipuzkoa, where Basque and Spanish are both everyday languages, so the cohort spans people who are close to monolingual in Spanish through to balanced Basque-Spanish bilinguals and speakers of three languages. The aim is to study how the language comprehension network varies with age of acquisition, exposure and proficiency. The data descriptor appeared in Scientific Data in 2026, and the release is on OpenNeuro as ds007111.

Composition

The paper and participants.tsv list 725 healthy adults aged 18 to 82 (389 women; mean age 34, median 27). Snapshot 1.1.1 holds 1,684 BOLD runs: 717 subjects have at least one Spanish run, 294 a Basque run and 57 an English run, and many subjects have repeated runs. Derivatives include CAT12 morphometry (atlas volumes, cortical surfaces and thickness) for 723 subjects, SPM12 probabilistic atlases of the auditory and visual comprehension contrasts, and a phenotype table with sociolinguistic, education and cognitive measures. The stimuli for all three languages are included.

Acquisition

All scans come from a 3 T Siemens Magnetom scanner. T1-weighted images are 3D MPRAGE at 1 mm isotropic with a 64-channel head coil. Functional runs are gradient-echo EPI; the two most common settings are TR 850 ms with 2.4 mm voxels and TR 2400 ms with 3 mm voxels. The task is an adaptation of the Pinel localizer: auditory and visual sentences for passive comprehension, mental arithmetic or a button press, plus flashing checkerboards.

Annotations

There are no image labels. Events files give onsets and conditions of every trial.

Known limitations

  • The original T1-weighted images are not on OpenNeuro. BCBL shares them only for justified requests under its own agreement, and asks researchers to use the CAT12 derivatives instead where possible.
  • Data come from many projects, so acquisition parameters, number of runs and languages tested differ between people. Some subjects have more than the 14 runs the paper states as the maximum.
  • Phenotype variables are missing for part of the cohort, with missing rates from 1 to 52 percent per variable.
  • Three subjects have T1-weighted scans of too low quality for morphometry.

Cohort

Aggregate numbers from the sources below. Bars are relative to the 725 subjects.

Contrast / sequence

Groups can overlap

  • BOLD fMRI 725 100%
  • T1-weighted 725 100%

From File tree of ds007111 snapshot 1.1.1 (CC0), NIfTI files counted per subject folder; Methods: Sample description and data harmonization in Quiñones et al. 2026, Unraveling the Complexity of Multilingual Comprehension, Scientific Data

Age

mean 34 ± 16.6, median 27, range 18 to 82

0
100
200
300
336
10-1920-2930-3940-4950-5960-6970-7980-89

age column in participants.tsv of ds007111 snapshot 1.1.1 (CC0), age and sex columns; Methods: Sample description and data harmonization in Quiñones et al. 2026, Unraveling the Complexity of Multilingual Comprehension, Scientific Data

Age by sex

Reported cross table. Missing cells were not published (fewer than 10 or not reported).

Female Male
  • 18
    70-79
    26
  • 31
    60-69
    18
  • 17
    50-59
    20
  • 26
    40-49
    28
  • 70
    30-39
    64
  • 184
    20-29
    152
  • 43
    10-19
    26

age and sex columns in participants.tsv of ds007111 snapshot 1.1.1 (CC0), age and sex columns

Contrast combinations

How many subjects have exactly each set of contrasts.

T1wboldSubjects with exactly this set
725

Methods: Sample description and data harmonization in Quiñones et al. 2026, Unraveling the Complexity of Multilingual Comprehension, Scientific Data

License and access

Our reading of the license, not legal advice. Before you use the data, read the original license and confirm that your use is allowed. We take no responsibility for how you use a dataset. Full disclaimer

Access
Open download

Download without an account

The BOLD runs, events files, stimuli, CAT12 structural derivatives, SPM12 group atlases and phenotype tables are an open download on OpenNeuro. The original T1-weighted images are not in the OpenNeuro release; BCBL shares them only for justified requests through a web form, after acceptance of a BCBL data use agreement.

Access page

Different parts of the data carry different licenses. The summary on the right shows the most restrictive answer per rule.

For OpenNeuro release ds007111 (snapshot 1.1.1), BOLD runs, derivatives, phenotype tables and stimuli.

Creative Commons Zero 1.0 Universal

Public domain dedication. Do anything with the data, including commercial use, without asking and without having to give credit.

dataset_description.json states "License" CC0. The README usage notes add that the data are for academic researchers, institutions and entities on condition of proper attribution, and that the authors' names may not be used to endorse commercial products without written permission.

Original license text Version read: 1.0 Checked 2026-10-07

What you can do

  • Yes
  • Yes
  • Yes
  • Yes
  • Yes

What you can share

  • Yes
  • Yes
  • Yes

What you must do

  • No
  • Share alike No
  • No
  • No
  • Manuscript review No
  • Release code No
  • Return results No
  • Delete after use No

Limits

  • No
  • Location limits No
For Original T1-weighted images, shared by BCBL on request only.

No published license or data use terms

The data holder has published no license and no data use terms. Any use, sharing or commercial right has to be agreed with the data holder, and nothing can be assumed to be allowed. Default copyright and data protection law still apply.

The BCBL page requires a scientific justification and acceptance of a BCBL data use agreement, whose text is not published.

Original license text Checked 2026-10-08

What you can do

  • Commercial use Not stated
  • Not stated
  • Train ML models Not stated
  • Create derived data Not stated
  • Publish results Not stated

What you can share

  • Share the data Not stated
  • Share derived data Not stated
  • Share trained models Not stated

What you must do

  • Cite or credit Not stated
  • Share alike Not stated
  • Sign an agreement Not stated
  • Ethics approval Not stated
  • Manuscript review Not stated
  • Release code Not stated
  • Return results Not stated
  • Delete after use Not stated

Limits

  • No re-identification Not stated
  • Location limits Not stated

Citation

Quiñones I, Carrión-Castillo A, Diez-Zabala I, et al. Unraveling the Complexity of Multilingual Comprehension: Neuroimaging and Linguistic Profiling in 700+ Adults. Scientific Data 13, 1315 (2026). doi:10.1038/s41597-026-07423-9

Sources

Every number on this page comes from one of these documents. Each chart names the table or page it is taken from. The raw numbers are in stats.csv.