Skip to content
Mammography Breast

VinDr-Mammo

VinDr-Mammo: A large-scale benchmark dataset for computer-aided diagnosis in full-field digital mammography

5,000 four-view full-field digital mammography exams (20,000 DICOM images) from two Hanoi hospitals, double read by radiologists with breast-level BI-RADS and density, bounding boxes for findings and a fixed train/test split.

Overview

VinDr-Mammo is a Vietnamese collection of full-field digital mammograms released by VinBigData on PhysioNet in 2022. It was built as a benchmark for computer-aided detection and diagnosis of breast findings and for predicting BI-RADS assessment and breast density, and it is one of the larger public digital mammography sets with radiologist annotations.

Composition

The dataset contains 5,000 exams with four images each: craniocaudal and mediolateral oblique views of both breasts, 20,000 DICOM images in total. The creators split the exams into 4,000 for training and 1,000 for testing with iterative stratification, so that BI-RADS categories, density levels and finding types have similar frequencies in both parts. Two CSV files hold the breast-level labels and the finding boxes, and a third keeps age and scanner model from the DICOM headers. The number of distinct women is not published.

Acquisition

Exams were sampled at random from the PACS of Hanoi Medical University Hospital and Hospital 108, covering 2018 to 2020, so screening and diagnostic exams are mixed. Images are "for presentation" mammograms from Siemens, IMS and Planmed units. Identifying text burned into image corners was blacked out, and only age and device information were kept in the headers.

Annotations

Three radiologists with 14 to 22 years of experience took part. Each exam was read independently by two of them, and a third, more senior reader settled disagreements. Each breast received a BI-RADS category (1 to 5) and a density category (A to D). Findings that needed follow-up (BI-RADS 3 or higher) were boxed and typed: mass, suspicious calcification, asymmetry (global or focal), architectural distortion, skin thickening or retraction, nipple retraction and suspicious lymph node. Benign BI-RADS 2 findings were not boxed.

Known limitations

  • No pathology confirmation; labels are radiologist consensus only.
  • Some finding types have fewer than 40 examples.
  • The authors note that the files are not fully DICOM-compliant.
  • Single country and two hospitals.

Cohort

Aggregate numbers from the sources below. Bars are relative to the largest value.

Anatomy

studies

  • Breast 5,000 100%

Split

images

  • Training 16,000 80%
  • Test 4,000 20%

View

images

  • Craniocaudal 10,000 50%
  • Mediolateral oblique 10,000 50%

License and access

Our reading of the license, not legal advice. Before you use the data, read the original license and confirm that your use is allowed. We take no responsibility for how you use a dataset. Full disclaimer

Access
Signed agreement

Sign a data use agreement, often reviewed by the provider

Access page

PhysioNet Restricted Health Data License 1.5.0

Research-only license for registered PhysioNet users who sign the matching data use agreement online. Unlike the credentialed license it asks for no identity check or human-subjects training. You must not share the data or try to identify people, and you must release the code behind your publications.

PhysioNet lists the PhysioNet Restricted Health Data License 1.5.0 and the matching data use agreement. The Scientific Data paper instead names the PhysioNet Credentialed Health Data License 1.5.0; the PhysioNet project page, which grants access, is taken as authoritative.

Original license text Version read: 1.5.0 Checked 2026-10-08

What you can do

  • Not stated
  • Not stated
  • Create derived data Not stated
  • Conditional

What you can share

  • No
  • Not stated
  • Share trained models Not stated

What you must do

  • Not stated
  • Share alike No
  • Yes
  • Ethics approval No
  • Manuscript review No
  • Yes
  • No
  • No

Limits

  • Yes
  • Location limits No

Citation

Nguyen HT, Nguyen HQ, Pham HH, et al. VinDr-Mammo: A large-scale benchmark dataset for computer-aided diagnosis in full-field digital mammography. Sci Data 10, 277 (2023). Pham HH, Nguyen Trung H, Nguyen HQ. VinDr-Mammo (version 1.0.0). PhysioNet (2022). https://doi.org/10.13026/br2v-7517

All numbers

Every number on this page, as stored in stats.csv, with its source.

MeasureBreakdownValueSource
Studiestotal
Exams; the number of distinct women is not reported
5,000nguyen2023
Methods: Data acquisition
Studiessplit=train 4,000nguyen2023
Methods: Data stratification
Studiessplit=test 1,000nguyen2023
Methods: Data stratification
Studiesanatomy=breast 5,000nguyen2023
Methods: Data acquisition
Imagestotal
Four images per exam
20,000nguyen2023
Methods: Data acquisition
Imagessplit=train
Four images per exam
16,000nguyen2023
Table 3 (8000 breasts x 2 views)
Imagessplit=test
Four images per exam
4,000nguyen2023
Table 3 (2000 breasts x 2 views)
Imagesview=CC
One CC and one MLO image per breast
10,000nguyen2023
Data Records
Imagesview=MLO
One CC and one MLO image per breast
10,000nguyen2023
Data Records

Sources

The keys used in the table above.