Contributing
Paste a link. Takes a minute, someone else writes the entry.
Open the formEvery dataset page has an Edit button that opens the file on GitHub.
Browse the repositoryThree files in one folder. The checker tells you what is missing.
Read the stepsThere are three ways to help, from least to most effort. An AI agent can do the third for you, see below.
- Suggest a dataset. Open an issue with the "Suggest a dataset" form and paste the link. Someone else writes the entry.
- Fix an entry. Every dataset page has an "Edit on GitHub" link. Change the file in the browser and open a pull request.
- Add a dataset. Follow the steps below.
Contribute with an AI agent
The repository ships an agent skill for contributions: contribute-imaging-dataset. Install it as described on the agent skills page, then ask the agent, for example, "Add BraTS 2023 to the Open Imaging Index". The skill makes the agent verify every fact online, look up vocabulary ids, run bun run validate and open a pull request. It asks you for what only you know, such as your GitHub handle. Agents that work in a clone without the skill find the same rules through AGENTS.md.
Pull requests written by agents are reviewed like any other. The person who asked the agent is responsible for the contribution.
Add a dataset
- Fork the repository and install Bun.
- Run
bun install, thenbun run new-dataset my-dataset-2024. This copiesdatasets/_template/todatasets/my-dataset-2024/. - Fill in
dataset.yaml. Editors with the YAML extension (VS Code: Red Hat YAML) autocomplete every field from the schema. - Write
README.mdin your own words. Do not copy the provider's text. - Add the numbers you can find to
stats.csv. A single total row is fine to start. See the data standard for what may go in. - If the dataset uses a license that has no file in
licenses/yet, add one and answer every rule with a quote. - Run
bun run validateand fix what it reports.bun run devshows the site with your entry. - Open a pull request. The checklist in the template covers what reviewers look at.
What reviewers check
- Every number has a source, and the source is public. Numbers computed from data held under a data use agreement are not accepted.
- License answers quote the license text. Silence stays
unspecified. - The README is original text, not copied from the dataset website or paper.
verified.dateis the day you checked the sources.
Licenses of this repository
- Code: MIT (LICENSE), for everything outside the data directories.
- Data: CC0 1.0 (LICENSE-DATA) for
datasets/,licenses/,vocab/anddocs/. Your contribution is published under these terms. Quotes from license texts and dataset citations belong to their authors.
The license summaries on this site help people find datasets. They are not legal advice. The license or agreement of each dataset is what counts.