OpenMind
The OpenMind Dataset: 3D head-and-neck MRI pooled from 800 OpenNeuro datasets for self-supervised learning
113,921 3D head and brain volumes (all MRI except 653 PET) from 800 OpenNeuro datasets and about 34,000 subjects, with diffusion-derived maps, defacing and anatomy masks and harmonized metadata, for self-supervised pre-training. CC BY 4.0 on Hugging Face.
Overview
OpenMind is a pre-training collection of 3D head and brain images assembled by the Division of Medical Image Computing at the German Cancer Research Center (DKFZ). The authors took every 3D MRI and every 4D diffusion scan they could find in 800 public OpenNeuro datasets, converted the diffusion scans into 3D maps, and released the result on Hugging Face together with a benchmark of 3D self-supervised learning methods, the code and the pre-trained models. It holds no task labels and is meant for pre-training, not for supervised training or evaluation.
Composition
The paper reports 113,921 volumes. Of these, 653 are PET, and 40,221 are fractional anisotropy, mean diffusivity and T2-weighted volumes computed from 13,407 diffusion scans. The volumes are grouped into 24 image types; T1w (42,732) and T2w (22,999) dominate, followed by FA and MD maps, FLAIR and MP2RAGE. Smaller groups include 784 SWI, 687 minimum intensity projections, 409 T2*-weighted and 76 T2* map volumes. The subject count differs inside the paper (34,139 in the text, 34,191 in the tables). Twelve source datasets contribute half of all volumes.
Acquisition
The data come from many independent studies, so protocols, scanners and resolutions vary. The harmonized metadata name Siemens, Philips and GE scanners and field strengths from 1.5 to 9.4 T, mostly 3 T, but manufacturer and field strength are missing for about a quarter of the images. Age, sex, handedness, BMI, race and health status are filled in only where the source dataset reported them.
Annotations
Each image has a defacing mask that marks anonymized regions and an anatomy mask that marks where tissue is present, created with an automated model where the source did not provide them. Two raters scored two example images per image type and source dataset for noise, blur and artifacts; the resulting image quality score from 1 (best) to 5 applies to all images of that type in that dataset.
Known limitations
- The quality score is per source dataset and image type, not per image.
- Many images are skull-stripped, defaced or face-blurred, which can affect reconstruction-based pre-training.
- Each source dataset keeps its own OpenNeuro license, and the same person may appear in more than one source dataset.
- Image types come from the source BIDS labels, which are named inconsistently; the authors say the grouping into 24 types is approximate.
Cohort
Aggregate numbers from the sources below. Bars are relative to the 34,139 subjects.
Sex
- Female 12,543 55%
- Male 10,300 45%
Covers 22,843 of 34,139 subjects.
Age
· range 2 to 100Modality
scans
- MRI 113,268 99%
- PET 653 <1%
Contrast / sequence
scans, values can overlap
- T1-weighted 42,732 38%
- T2-weighted 22,999 20%
- FLAIR 5,583 5%
- MP2RAGE 2,859 3%
- ADC map 1,757 2%
- Susceptibility-weighted 784 <1%
- T1 map 557 <1%
- MR angiography 488 <1%
- Proton density weighted 473 <1%
- T2*-weighted 409 <1%
- T2 map 106 <1%
- Diffusion-weighted 55 <1%
Condition
subjects, values can overlap
- Healthy control 5,262 15%
Scanner vendor
scans
- Siemens Healthineers 64,415 57%
- Philips 14,580 13%
- GE HealthCare 8,114 7%
Field strength
scans
- 3 T 75,107 66%
- 1.5 T 5,948 5%
- 7 T 1,329 1%
- 9.4 T 86 <1%
License and access
Our reading of the license, not legal advice. Before you use the data, read the original license and confirm that your use is allowed. We take no responsibility for how you use a dataset. Full disclaimer
Download without an account
Creative Commons Attribution 4.0 International
Use, share and adapt the data for any purpose, including commercial use, as long as you credit the creators.
The Hugging Face dataset card declares cc-by-4.0, and the paper (Table 1) lists the dataset as CC-BY-4.0. The images come from OpenNeuro datasets that each carry their own license.
What you can do
- Yes
- Yes
- Yes
- Yes
What you can share
- Yes
- Yes
- Yes
What you must do
- Yes
- Share alike No
- No
- No
- Manuscript review No
- Release code No
- Return results No
- Delete after use No
Limits
- No
- Location limits No
Citation
Wald T, Ulrich C, Suprijadi J, Ziegler S, Nohel M, Peretzke R, Köhler G, Maier-Hein KH. An OpenMind for 3D medical vision self-supervised learning. arXiv:2412.17041 (2024).
@article{wald2024openmind,
title={An OpenMind for 3D medical vision self-supervised learning},
author={Wald, Tassilo and Ulrich, Constantin and Suprijadi, Jonathan and Ziegler, Sebastian and Nohel, Michal and Peretzke, Robin and K{\"o}hler, Gregor and Maier-Hein, Klaus H.},
journal={arXiv preprint arXiv:2412.17041},
year={2024}
} All numbers
Every number on this page, as stored in stats.csv, with its source.
| Measure | Breakdown | Value | Source |
|---|---|---|---|
| Subjects | total Table 1 and Appendix B.3 give 34,191; the metadata file has 34,139 unique dataset and subject pairs. The same person can appear in more than one OpenNeuro dataset. | 34,139 | wald2024 Section 2.1 |
| Scans | total 3D volumes, including 653 PET and 40,221 volumes derived from 13,407 diffusion scans; the metadata file has 113,921 rows | 113,921 | wald2024 Appendix B.3 |
| Scans | modality=PT footnote 1 | 653 | wald2024 Section 2.1 |
| Scans | modality=MR all rows except PET | 113,268 | openmind-metadata modality |
| Scans | contrast=T1w | 42,732 | wald2024 Table 9 |
| Scans | contrast=T2w includes 13,407 T2-weighted volumes derived from diffusion scans (metadata column derived_from) | 22,999 | wald2024 Table 9 |
| Scans | contrast=FLAIR | 5,583 | wald2024 Table 9 |
| Scans | contrast=MP2RAGE UNIT1 (927) and UNIT1_denoised (724) are listed separately | 2,859 | wald2024 Table 9 |
| Scans | contrast=ADC | 1,757 | wald2024 Table 9 |
| Scans | contrast=swi minimum intensity projections (minIP, 687) are listed separately | 784 | wald2024 Table 9 |
| Scans | contrast=T1map | 557 | wald2024 Table 9 |
| Scans | contrast=angio | 488 | wald2024 Table 9 |
| Scans | contrast=PDw | 473 | wald2024 Table 9 |
| Scans | contrast=T2starw T2* maps (T2starmap, 76) are listed separately | 409 | wald2024 Table 9 |
| Scans | contrast=T2map | 106 | wald2024 Table 9 |
| Scans | contrast=dwi remaining DWI volumes; the 13,407 diffusion scans were converted into FA, MD and T2w volumes | 55 | wald2024 Table 9 |
| Scans | vendor=siemens | 64,415 | openmind-metadata manufacturer |
| Scans | vendor=philips | 14,580 | openmind-metadata manufacturer |
| Scans | vendor=ge 26,811 rows have no manufacturer | 8,114 | openmind-metadata manufacturer |
| Scans | field_strength=1.5 | 5,948 | openmind-metadata magnetic_field_strength |
| Scans | field_strength=3 75,069 rows with 3.0 plus 38 with 2.89362 (the nominal value some 3 T Siemens scanners report) | 75,107 | openmind-metadata magnetic_field_strength |
| Scans | field_strength=7 335 rows with 7.0 plus 994 with 6.98 | 1,329 | openmind-metadata magnetic_field_strength |
| Scans | field_strength=9.4 30,104 rows have no field strength; 1,342 rows with 3.95064 and 5 with 15000.0 are left out | 86 | openmind-metadata magnetic_field_strength |
| Subjects | sex=female subjects whose images list female | 12,543 | openmind-metadata unique_id, sex |
| Subjects | sex=male subjects whose images list male; 11,296 subjects have no sex | 10,300 | openmind-metadata unique_id, sex |
| Subjects | condition=healthy 3,285 subjects are listed as ill, the rest have no health status | 5,262 | openmind-metadata unique_id, health_status |
| Subjects | age=0-9 age as given in the harmonized metadata; 20,951 subjects have an age | 591 | openmind-metadata unique_id, age |
| Subjects | age=10-19 | 3,060 | openmind-metadata unique_id, age |
| Subjects | age=20-29 | 11,061 | openmind-metadata unique_id, age |
| Subjects | age=30-39 | 2,437 | openmind-metadata unique_id, age |
| Subjects | age=40-49 | 971 | openmind-metadata unique_id, age |
| Subjects | age=50-59 | 910 | openmind-metadata unique_id, age |
| Subjects | age=60-69 | 983 | openmind-metadata unique_id, age |
| Subjects | age=70-79 | 676 | openmind-metadata unique_id, age |
| Subjects | age=80-89 | 257 | openmind-metadata unique_id, age |
| Subjects | age=90+ | 5 | openmind-metadata unique_id, age |
| Median age | total 20,951 subjects with an age | 24 | openmind-metadata unique_id, age |
| Minimum age | total | 2 | openmind-metadata unique_id, age |
| Maximum age | total | 100 | openmind-metadata unique_id, age |
Sources
The keys used in the table above.
- wald2024 Wald et al. 2024, An OpenMind for 3D medical vision self-supervised learning, arXiv:2412.17041v2 paper
- hf-openmind-card Hugging Face dataset card MIC-DKFZ/OpenMind website
- openmind-metadata openneuro_metadata.csv of MIC-DKFZ/OpenMind (CC BY 4.0, no login), one row per image; subjects are unique dataset and subject pairs from unique_id computed