Skip to content

About

Finding a medical imaging dataset usually means reading papers, websites and data use agreements one by one. Lists of links exist, but they rarely say how many subjects have the contrasts you need, or whether you may train a commercial model on the data. Open Imaging Index collects that information in one place, in a form that can be searched and combined.

What the index publishes

  • A description of each dataset in our own words, with a link to its home.
  • Aggregate numbers about the cohort: counts per contrast, sex, age bin, condition and more. Never data about individual subjects.
  • A breakdown of each license into the same set of questions, with the sentence each answer is based on.

Where the numbers come from

Numbers come from the dataset's own documentation and papers, or are counted from metadata that is itself openly licensed. Nothing is computed from data that was obtained under a data use agreement. Small counts for combinations of attributes that we compute ourselves are not published, because they could identify people. The data standard has the details.

Reading the license summaries

The summaries help you compare datasets and find candidates. They are not legal advice. Licenses change, and a summary can be wrong. Before you use a dataset, read its license or agreement: that text is what counts. Each entry shows the date it was last checked. If you find a mistake, use the Report button on the page. The disclaimer has the details.

Open source

The code is MIT licensed (LICENSE). The data, meaning dataset entries, license breakdowns, vocabularies and docs, is CC0 (LICENSE-DATA). Everything lives in one GitHub repository, and the whole catalog is available as catalog.json for scripts and other tools.