Skip to content

Agent skills

Let your AI assistant search the index for you, or add a dataset it is missing. Two skills in one plugin, in the open SKILL.md format that Claude and most other agent tools read.

Find datasets

find-imaging-datasets

Turns a need into a ranked shortlist: matching datasets, subject counts for your cohort, access and the license terms that matter for your use, checked against the official pages.

Try asking
  • Which brain MRI datasets have T1w and FLAIR for at least 500 subjects?
  • Find chest X-ray datasets I may use to train a commercial model.
  • How many women aged 60 to 80 with FLAIR scans are in open datasets?

Works anywhere the agent can fetch a web page.

Add or correct a dataset

contribute-imaging-dataset

Adds a missing dataset, or fixes an entry, as a pull request. Strict rules: every fact verified online in a public source, nothing computed from gated data, license answers quoted, and the validator must pass.

Try asking
  • Add BraTS 2023 to the Open Imaging Index.
  • The subject count of IXI looks wrong. Check it and open a correction.
  • Break down the license of the KiTS23 dataset.

Needs a shell with git, gh and Bun, e.g. Claude Code or Codex.

Install

Add the marketplace and install the plugin. Run this inside Claude Code:

/plugin marketplace add Spenhouet/open-imaging-index
/plugin install open-imaging-index@open-imaging-index

Or from your shell:

claude plugin marketplace add Spenhouet/open-imaging-index
claude plugin install open-imaging-index@open-imaging-index

Claude picks the right skill from your request. You can also call one directly with /open-imaging-index:find-imaging-datasets or /open-imaging-index:contribute-imaging-dataset. Updates arrive with /plugin marketplace update open-imaging-index.

Agents without skills can read llms.txt and the full catalog at catalog.json.

Full skill find-imaging-datasets/SKILL.md

Find medical imaging datasets

The Open Imaging Index (https://spenhouet.com/open-imaging-index/) lists medical imaging datasets with aggregate cohort numbers and a license breakdown into the same yes/no rules for every dataset. This skill turns a person's need into a ranked, verified shortlist.

Rules

  1. Work from the published catalog, then verify. The catalog says what the index recorded on its verified.date. Before you recommend a dataset for a decision (a grant, a product, a training run), open its homepage and license page and confirm the facts that decision depends on. Say which facts you confirmed online and which you only took from the index.
  2. Never present a license summary as legal advice. Whenever your answer mentions what a license allows, include this sentence: "License answers are the Open Imaging Index's interpretation, not legal advice. Read the original license and confirm your use yourself (https://spenhouet.com/open-imaging-index/disclaimer/)." Every answer in the index links to the license text. For commercial use or model sharing, quote the license sentence and tell the person to read the full license.
  3. Never invent numbers. Use the numbers in the catalog or numbers you read in a source in this session. If a count is a range (see "Cohort numbers"), give the range, not a guess inside it.
  4. Never download data that needs registration or an agreement on the person's behalf, and never ask them for credentials.
  5. If nothing fits, say so. Do not stretch a weak match into a recommendation.

Steps

1. Pin down the need

Ask only for what is missing and matters. Typical questions:

  • Modality and contrasts, e.g. MRI with T1w and FLAIR, contrast CT, chest X-ray.
  • Anatomy and condition, e.g. brain, glioma; chest, pneumothorax.
  • Cohort: minimum subjects, age range, sex, healthy controls needed.
  • Labels or task: segmentation masks, classification labels, reports.
  • Use of the data: commercial product, training models, publishing model weights, re-sharing data.
  • Access the person can handle: open download, free registration, signed agreement, credentialed access (e.g. PhysioNet with CITI training), application review, paid.

If the person gives a broad request ("brain MRI datasets"), proceed with sensible defaults and state them.

2. Load the catalog

Download the full catalog (CC0, about 250 KB):

curl -sL https://spenhouet.com/open-imaging-index/catalog.json -o catalog.json

If you cannot run shell commands, fetch the same URL with your web tool. Structure:

  • datasets[]: one object per dataset.
    • id: page at https://spenhouet.com/open-imaging-index/datasets/<id>/.
    • meta: the dataset.yaml content: name, full_name, summary, homepage, doi, year, modalities, contrasts, anatomy, conditions, tasks, formats, countries, access.type, access.url, citation, sources, verified.date.
    • facets: values per dimension (modality, contrast, anatomy, condition, vendor, field_strength, country, ...), merged from metadata and stats.
    • totals: totals per measure (subjects, studies, scans, images, slides).
    • stats[]: every number, as {measure, by, value, approx}. by maps dimensions to values, e.g. {"sex": "female", "age": "60-69"}. An empty by is the total.
    • rules: the combined license answer per rule (yes, no, conditional, unspecified).
    • licenses[]: license ids with applies_to when a dataset mixes licenses.
  • licenses[]: each license file with rules.<rule>.value, quote, note, url, summary.
  • vocab: vocabularies with labels, synonyms and parent terms (condition glioblastoma has parent glioma).

Vocabulary values are ids, not free text: modalities are DICOM codes (MR, CT, PT, DX for X-ray, MG, US, SM for pathology slides, OP for fundus), MR contrasts are BIDS suffixes (T1w, T1w_ce, T2w, FLAIR, dwi, bold). Map the person's words to ids with vocab.terms.<dimension> labels and synonyms, and include child terms of a condition.

3. Filter and count

Match on facets, rules and meta.access.type. For license needs, check rules:

Need Rule Acceptable answer
Commercial use commercial_use yes (report conditional separately)
Train models model_training yes or conditional with the condition stated
Publish or sell trained models share_model_weights yes
Re-host or share the data redistribute_original yes
No agreement to sign signed_agreement no
No ethics approval needed ethics_approval no

unspecified means the license text is silent. Treat it as "ask the provider", never as yes.

Cohort numbers

For cohort questions, count from stats with measure subjects:

  • A row whose by matches the question exactly gives an exact count.
  • contrast_set rows give exact counts per combination of contrasts. Subjects with both T1w and FLAIR are the sum of all sets that contain both.
  • With only single-dimension rows, give bounds. At least max(0, a + b - N) and at most min(a, b), where a and b are the two counts and N is the total.
  • Age bins use completed years: 60-69 covers ages 60 up to the 70th birthday. A bin that only partly overlaps the requested range adds to the upper bound only.
  • Values written with approx: true are approximate in the source.

The website does the same arithmetic. Give the person a link that reproduces the filter, for example:

https://spenhouet.com/open-imaging-index/?contrasts=FLAIR,T1w&sex=female&age=60-80&min=200&rules=commercial_use

Query parameters: q (search), modality, anatomy, condition, task, access, license, format, vendor, field_strength, country (comma-separated ids), contrasts (all required), sex, age (from-to in years, 100 means no upper limit), min (minimum matching subjects), rules (comma-separated rule ids that must be in the user's favor), lenient=1 (count conditional answers as allowed).

4. Verify online

For each dataset you put on the shortlist:

  1. Open meta.homepage and confirm the dataset still exists and how to get it.
  2. If the person's use depends on a license rule, open the license url, find the sentence in the quote, and confirm it still says that.
  3. If verified.date is more than a year old, or what you read differs from the index, say so plainly and suggest a correction (see below).

5. Answer

Give a ranked shortlist, best fit first. For each dataset:

  • Name, linked to its page in the index, plus the official homepage.
  • Why it fits, with the numbers that matter (subjects, matching subjects or bounds, contrasts, labels).
  • Access type and the key license answers for this person's use, with the deciding quote where it matters.
  • What you verified online and what you did not.

End with the gaps: requirements no dataset meets, and anything the person must check themselves (license text, data quality, ethics).

When the index is missing something

If you know or find a dataset that fits but is not in the index, or an entry is wrong:

Full skill contribute-imaging-dataset/SKILL.md

Contribute a dataset to the Open Imaging Index

The index lives in https://github.com/Spenhouet/open-imaging-index. Each dataset is one folder datasets/<id>/ with dataset.yaml, README.md and stats.csv. Licenses live in licenses/, allowed values in vocab/. The site is built from these files, and bun run validate checks them in CI.

People rely on these entries to choose datasets and to judge what they may legally do with them. A wrong number or a wrong license answer does real damage. Accuracy beats completeness every time.

Hard rules

You must follow every rule. If a rule cannot be met, stop and tell the person why, instead of working around it.

  1. Verify everything online, in this session. Every number, date, license answer and description detail must come from a public source you opened during this task: the dataset's paper or data descriptor, its official website, its license text, or an openly downloadable metadata file. Your memory, other catalogs, blog posts and earlier versions of this index are not sources. When a source cannot be fetched, the fact stays out.
  2. Never invent or estimate. No number unless a source states it, or it is counted from an open file. Write approximate values exactly as the source does (~600 for "nearly 600"). Sum parts into a total only when the parts are disjoint and complete, and say so in the note.
  3. Every stats row names its source and where in it. The source column is a key under sources in dataset.yaml. The key names one document (baid2021, tcia-brats2021), never a kind (paper, website). The where column gives the table, figure, page or section, e.g. Table 2 or p. 4.
  4. Never compute from data behind a gate. Counting from files is allowed only when anyone can download them without an account, registration or agreement. Never use data the person or you obtained under a data use agreement, credentialed access (PhysioNet, ADNI, UK Biobank) or competition rules, not even for counts. For counts you compute, leave out every cell that combines two or more dimensions and is below 10. The validator enforces this for computed sources.
  5. Write the README in your own words. Do not copy sentences from the website or paper.
  6. License answers quote the license. Every answer in a licenses/*.yaml file other than a plain "no duty" quotes the exact sentence it rests on (quote), with source if the quote is not from the license text itself. If the text is silent on a permission, the answer is unspecified, never yes or no. A duty or limit is no only when nothing in the text asks for it. Any clause that restricts purpose, users, location or time, or requires an action, makes that rule yes or conditional with the quote.
  7. Vocabulary ids are looked up, never guessed. Run bun run lookup "<term>" mondo (or uberon, hp) and copy the id. Add a term only when no existing term fits. Never rename or delete existing terms.
  8. Run the checks yourself and read the output. bun run validate must report 0 errors before you open a pull request. Fix warnings about your files where a source allows it. If you changed anything outside datasets/, licenses/ and vocab/, also run bun run check, bun run lint and bunx vitest --run.
  9. Do not guess the person's details. Ask for anything only they can provide, and wait for the answer.
  10. One dataset per pull request. Keep the diff to what the task needs.

Ask the person

Before writing files, get these from the person if they are not already clear:

  • Which dataset, and which release or version. If the name is ambiguous (BraTS 2021 vs 2023, MIMIC-CXR vs MIMIC-CXR-JPG), confirm.
  • Their GitHub handle, for verified.by. If they prefer, use <handle>-agent.
  • Whether they hold access under an agreement. If yes, remind them that nothing derived from that access may go into the index.
  • For corrections: what is wrong, and the source that shows the right value, if they have one.
  • Permission to fork the repository, push a branch and open a pull request in their name.

Set up

gh repo fork Spenhouet/open-imaging-index --clone   # or: git clone https://github.com/Spenhouet/open-imaging-index
cd open-imaging-index
bun install                                         # needs Bun, and Node 22+ for building the site
git switch -c add-<dataset-id>                      # or fix-<dataset-id>-<topic>

Read these files completely before you change anything:

  • docs/standard.md: the data standard. It is the reference for every field.
  • datasets/_template/: the starting point.
  • One finished entry close to your dataset, e.g. datasets/ixi/ (open data with computed counts), datasets/brats-2021/ (two alternative licenses) or datasets/nih-chestxray14/ (computed cross tables).
  • vocab/license-rules.yaml and one custom license file such as licenses/LicenseRef-OASIS-DUA.yaml, if a license file is involved.

Add a dataset

  1. Check it is not already tracked. Search datasets/*/dataset.yaml for the name, full name, DOI and homepage domain (grep -ril "<term>" datasets). Also look at https://spenhouet.com/open-imaging-index/?q=. If it exists, switch to "Edit an entry".
  2. Collect sources. Find the data descriptor paper (DOI), the official website, the download or access page and the license or agreement text. Prefer the paper for numbers and the provider's own page for access and license.
  3. Scaffold: bun run new-dataset <id> with a lowercase id like brats-2021.
  4. Fill dataset.yaml. Use only vocabulary ids (vocab/*.yaml). summary is 50 to 320 characters. year is the first public release. access.type is one of open, registration, signed_agreement, credentialed, application, paid. With several licenses, give each an applies_to, and set license_combine: any only when the same data is offered under alternative licenses. Set verified.date to today's date (date +%F).
  5. Licenses. If the dataset uses a license that already has a file in licenses/, reference it. Otherwise create licenses/<SPDX id>.yaml, or licenses/LicenseRef-<Name>.yaml for a custom agreement, answering every rule in vocab/license-rules.yaml under the hard rules above. Record the version or date of the text you read, and commercial_license if the provider sells one.
  6. Write stats.csv. Header: measure,by,value,source,where,note. Add every number the sources give: totals, per contrast, the exact contrast combinations (contrast_set), sex, age bins (completed years, 60-69), conditions, scanners, field strengths, countries, splits, and cross tables where reported. Use the right measure: subjects for people, studies for sessions, scans for volumes, images for 2D images, slides for whole-slide images.
  7. Write README.md. 150 to 400 words in your own words: one overview paragraph, then ## Composition, ## Acquisition, ## Annotations (if any) and ## Known limitations. Facts only.
  8. Validate: bun run validate. Fix every error and re-run until it reports 0 errors. Then look at the page: bun run dev and open /datasets/<id>/.
  9. Open the pull request (see below).

Edit an entry

  1. Open the files of the entry and find the claim in question.
  2. Verify the correct value online, in a source you open now. If the source agrees with the index, report that and stop.
  3. Change only what the source supports. Update source and where for every row you touch. Keep notes accurate.
  4. Update verified.date and verified.by only if you re-checked the whole entry against its sources. Otherwise leave them and say in the pull request what you checked.
  5. Run bun run validate until it reports 0 errors.

For a license change, edit the file in licenses/ and update version and verified. Every dataset that uses it changes with it, so name them in the pull request (grep -rl "license: <id>" datasets).

Open the pull request

Commit with a message that names the dataset and what changed, push, and open the pull request with gh pr create. Fill in the checklist from .github/pull_request_template.md honestly. In the description, list:

  • The sources you used, with links.
  • Anything you could not verify and therefore left out.
  • Any judgment call (an ambiguous license sentence, a total you summed, a value where sources disagree).

Report back

Tell the person what you added or changed, the pull request link, what you left out and why, and the validator result. Never claim a check passed unless you ran it and saw it pass.