Narratives (ds002345)
The Narratives fMRI dataset for evaluating models of naturalistic language comprehension
Story-listening fMRI of 345 healthy adults from Princeton: 891 BOLD runs across 27 spoken stories, with T1w (some T2w) anatomy, audio stimuli and time-stamped transcripts, in BIDS under CC0.
Overview
Narratives pools functional MRI recorded while adults listened to spoken stories in the Hasson and Norman labs at the Princeton Neuroscience Institute between October 2011 and September 2018. The authors shared it on OpenNeuro as ds002345 under a CC0 waiver and described it in Scientific Data in 2021. It serves as a benchmark for models of language and narrative comprehension and for naturalistic analyses such as intersubject correlation.
Composition
Snapshot 1.1.4 covers 345 adults aged 18 to 53 (mean 22.2 years, 204 reporting female). Subjects heard one or more of 27 stories lasting from about 3 to 56 minutes, about 4.6 hours of unique audio in total, and the paper counts 891 functional scans. Subjects who took part in several sub-studies keep one identifier, and participants.tsv lists the stories, experimental conditions and comprehension scores per subject. Every subject has a T1-weighted image, and 46 also have a T2-weighted image. The release includes the audio files, plain transcripts and word- and phoneme-level time stamps.
Acquisition
All scans come from the Scully Center for Neuroimaging in Princeton with a 1.5 s repetition time. The earlier sub-studies used a 3 T Siemens Skyra with a 20-channel head coil and 3 x 3 x 4 mm EPI. Later ones used a 3 T Siemens Prisma with a 64-channel coil and multiband EPI at 2 or 2.5 mm, and T2-weighted images were only acquired in the 2018 sub-studies. Images were converted with dcm2niix, defaced with pydeface and organized in BIDS 1.2.1.
Annotations
There are no image labels. Annotations are on the stimulus side: transcripts aligned in time to the audio. The full DataLad distribution adds fMRIPrep outputs, AFNI-postprocessed time series and MRIQC quality metrics.
Known limitations
- The sample is mostly young adults from a university community, and both native and non-native English speakers are included.
- Acquisition parameters differ between sub-studies, so data from the Skyra and Prisma groups need harmonization.
- The authors recommend excluding specific scans, flagged by their intersubject correlation checks and listed in code/scan_exclude.json.
- The paper counts 891 functional scans, while the snapshot 1.1.4 file listing holds 873 BOLD NIfTI files.
- Ages in participants.tsv are given per story, so a subject scanned across years has several ages.
Cohort
Aggregate numbers from the sources below. Bars are relative to the 345 subjects.
Sex
- Female 204 59%
- Male 141 41%
Age
mean 22.2± 5.1 · range 18 to 53No age bins reported.
Contrast combinations
How many subjects have exactly each set of contrasts.
| T1w | T2w | bold | Subjects with exactly this set |
|---|---|---|---|
299 | |||
46 |
Contrast / sequence
subjects, values can overlap
- BOLD fMRI 345 100%
- T1-weighted 345 100%
- T2-weighted 46 13%
Condition
subjects, values can overlap
- Healthy control 345 100%
Scanner vendor
subjects
- Siemens Healthineers 345 100%
Field strength
subjects
- 3 T 345 100%
Country
subjects
- United States 345 100%
License and access
Our reading of the license, not legal advice. Before you use the data, read the original license and confirm that your use is allowed. We take no responsibility for how you use a dataset. Full disclaimer
Download without an account
Creative Commons Zero 1.0 Universal
Public domain dedication. Do anything with the data, including commercial use, without asking and without having to give credit.
dataset_description.json of snapshot 1.1.4 states "License" CC0; the paper states that all MRI data and metadata are released under CC0.
What you can do
- Yes
- Yes
- Yes
- Yes
What you can share
- Yes
- Yes
- Yes
What you must do
- No
- Share alike No
- No
- No
- Manuscript review No
- Release code No
- Return results No
- Delete after use No
Limits
- No
- Location limits No
Citation
Nastase SA, Liu YF, Hillman H, et al. The "Narratives" fMRI dataset for evaluating models of naturalistic language comprehension. Scientific Data 8, 250 (2021). doi:10.1038/s41597-021-01033-3
All numbers
Every number on this page, as stored in stats.csv, with its source.
| Measure | Breakdown | Value | Source |
|---|---|---|---|
| Subjects | total participants.tsv and the OpenNeuro snapshot summary also list 345 subjects | 345 | nastase2021 Methods: Participants |
| Subjects | condition=healthy all subjects reported normal hearing and no history of neurological disorders | 345 | nastase2021 Methods: Participants |
| Subjects | sex=female reported female; participants.tsv also gives 204 F | 204 | nastase2021 Methods: Participants |
| Subjects | sex=male sex M | 141 | ds002345-participants sex column |
| Subjects | contrast=bold | 345 | ds002345-files sub-*/func/*_bold.nii.gz |
| Subjects | contrast=T1w | 345 | ds002345-files sub-*/anat/*_T1w.nii.gz |
| Subjects | contrast=T2w | 46 | ds002345-files sub-*/anat/*_T2w.nii.gz |
| Subjects | contrast_set=bold+T1w+T2w | 46 | ds002345-files sub-*/anat and sub-*/func |
| Subjects | contrast_set=bold+T1w | 299 | ds002345-files sub-*/anat and sub-*/func |
| Subjects | field_strength=3 3 T Siemens Magnetom Skyra and Prisma | 345 | nastase2021 Methods: MRI data acquisition |
| Subjects | vendor=siemens | 345 | nastase2021 Methods: MRI data acquisition |
| Subjects | country=US recruited in Princeton, NJ | 345 | nastase2021 Methods: Participants |
| Scans | total sum of 873 BOLD, 378 T1w and 47 T2w files, which are disjoint and are all NIfTI files under anat and func | 1,298 | ds002345-files sub-*/anat and sub-*/func NIfTI files |
| Scans | contrast=bold the paper reports 891 functional scans (Abstract) | 873 | ds002345-files sub-*/func/*_bold.nii.gz |
| Scans | contrast=T1w | 378 | ds002345-files sub-*/anat/*_T1w.nii.gz |
| Scans | contrast=T2w | 47 | ds002345-files sub-*/anat/*_T2w.nii.gz |
| Mean age | total | 22.2 | nastase2021 Methods: Participants |
| Age SD | total | 5.1 | nastase2021 Methods: Participants |
| Minimum age | total | 18 | nastase2021 Methods: Participants |
| Maximum age | total | 53 | nastase2021 Methods: Participants |
Sources
The keys used in the table above.
- nastase2021 Nastase et al. 2021, The "Narratives" fMRI dataset for evaluating models of naturalistic language comprehension, Scientific Data paper
- openneuro-ds002345-v114 OpenNeuro ds002345 snapshot 1.1.4 (dataset_description.json, CHANGES and snapshot summary) website
- ds002345-participants participants.tsv of ds002345 1.1.4 (CC0), counted per row from the sex column computed
- ds002345-files File listing of ds002345 1.1.4 (CC0), counted from sub-*/anat and sub-*/func NIfTI files computed