[Dataset] Add AOMIC-PIOP1 (ds002785) - Amsterdam Multimodal MRI (216 subjects)
#12 opened on Dec 9, 2025
Repository metrics
- Stars
- (1 star)
- PR merge metrics
- (PR metrics pending)
Description
Dataset Info
| Field | Value |
|---|---|
| Name | AOMIC-PIOP1 (Population Imaging of Psychology 1) |
| Source | OpenNeuro ds002785 |
| Paper | Snoek et al., Scientific Data 2021 |
| License | CC0 (Public Domain) |
| Subjects | 216 |
| Format | BIDS |
| HuggingFace Target | hugging-science/aomic-piop1 |
Description
AOMIC-PIOP1 is part of the Amsterdam Open MRI Collection - multimodal 3T MRI with detailed demographics and psychometric variables. This is the smallest of the 3 AOMIC datasets, making it ideal as a "tracer bullet" before tackling the larger ones.
Data includes:
- T1-weighted structural MRI
- Diffusion-weighted MRI
- Resting-state fMRI
- Task-based fMRI (emotion, working memory, faces, etc.)
- Physiological recordings (cardiac, respiratory)
- Demographics + psychometrics
Why This Matters
- Tracer bullet - Smallest AOMIC, test the pattern before ID1000 (928 subjects)
- Rich multimodal - Structural + diffusion + fMRI
- Task fMRI - Valuable for cognitive neuroscience
- Proves pipeline - Tests Sequence(Nifti()) for multiple runs
Exact Schema
from datasets import Features, Value
from datasets.features import Nifti, Sequence
def get_aomic_piop1_features() -> Features:
"""AOMIC-PIOP1 schema - one row per SUBJECT.
Note: Following arc.py pattern, we use a single `bold` list for ALL
functional runs (rest + tasks). Task info is in the filename (e.g.,
_task-rest_, _task-workingmemory_). Separation can be done downstream.
"""
return Features({
"subject_id": Value("string"),
# Structural
"t1w": Nifti(),
# Diffusion
"dwi": Sequence(Nifti()), # May have multiple runs
# Functional (all runs: rest + tasks)
"bold": Sequence(Nifti()), # All BOLD runs (*_bold.nii.gz)
# Metadata from participants.tsv
"age": Value("float32"),
"sex": Value("string"),
"handedness": Value("string"), # Verify column name (hand vs handedness)
})
Implementation Note: Verify
participants.tsvcolumn names. BIDS useshandednessbut some datasets usehand. Check actual file during implementation.
Directory Structure
ds002785/
├── participants.tsv
├── dataset_description.json
└── sub-XXXX/
├── anat/
│ └── sub-XXXX_T1w.nii.gz
├── dwi/
│ ├── sub-XXXX_dwi.nii.gz
│ ├── sub-XXXX_dwi.bval
│ └── sub-XXXX_dwi.bvec
└── func/
├── sub-XXXX_task-rest_bold.nii.gz
├── sub-XXXX_task-workingmemory_bold.nii.gz
├── sub-XXXX_task-faces_bold.nii.gz
└── ... (multiple tasks)
Files to Create
src/bids_hub/datasets/aomic_piop1.py # Dataset module
src/bids_hub/validation/aomic.py # Shared AOMIC validation
scripts/download_aomic_piop1.sh # Download script
tests/test_aomic_piop1.py # Tests
docs/dataset-cards/aomic-piop1.md # Dataset card (follow arc-aphasia-bids.md pattern)
Implementation Steps
-
Create download script (
scripts/download_aomic_piop1.sh)aws s3 sync --no-sign-request s3://openneuro.org/ds002785 "$TARGET_DIR" -
Create dataset module (
src/bids_hub/datasets/aomic_piop1.py)- Implement
build_aomic_piop1_file_table()- walk sub-*/anat/, dwi/, func/ - Implement
get_aomic_piop1_features()- schema above - Use
find_all_niftis()for Sequence features (multiple runs)
- Implement
-
Add CLI commands (in
cli.py)@aomic.command() def piop1(): """AOMIC-PIOP1 dataset commands.""" -
Add validation (
src/bids_hub/validation/aomic.py)- Check participants.tsv exists
- Check T1w exists per subject
- Report available modalities
-
Upload to HuggingFace
uv run bids-hub aomic piop1 build /path/to/ds002785 \ --hf-repo hugging-science/aomic-piop1 --no-dry-run- Use
num_shards=216(one per subject)
- Use
Acceptance Criteria
- Download script works
-
uv run bids-hub aomic piop1 validate <path>passes -
uv run bids-hub aomic piop1 build <path> --dry-runsucceeds - Tests pass
- Dataset uploaded to
hugging-science/aomic-piop1 - HuggingFace README.md with proper frontmatter, usage examples, and citation
-
docs/dataset-cards/aomic-piop1.mdadded (followarc-aphasia-bids.mdpattern)
Resources
Citation
@article{snoek2021aomic,
title={The Amsterdam Open MRI Collection, a set of multimodal MRI datasets for individual difference analyses},
author={Snoek, Lukas and van der Miesen, Maite M and others},
journal={Scientific Data},
volume={8},
number={1},
pages={85},
year={2021}
}
Related Issues
After this is complete, tackle:
- AOMIC-ID1000 (ds003097) - 928 subjects
- AOMIC-PIOP2 (ds002790) - 226 subjects