NM000402: ieeg dataset, 29 subjects#
Multi-Expert Seizure Annotation
Access recordings and metadata through EEGDash.
Citation: William K. Ojemann, Daniel Zhou, Caren Armstrong, Nishant Sinha, Brian Litt, Erin Conrad (2019). Multi-Expert Seizure Annotation. 10.82901/nemar.nm000402
Modality: ieeg Subjects: 29 Recordings: 78 License: CC-BY-4.0 Source: nemar
Metadata: Complete (100%)
29-participant iEEG dataset — Multi-Expert Seizure Annotation.
Quickstart#
Install
pip install eegdash
Access the data
from eegdash.dataset import NM000402
dataset = NM000402(cache_dir="./data")
# Get the raw object of the first recording
raw = dataset.datasets[0].raw
print(raw.info)
Filter by subject
dataset = NM000402(cache_dir="./data", subject="01")
Advanced query
dataset = NM000402(
cache_dir="./data",
query={"subject": {"$in": ["01", "02"]}},
)
Iterate recordings
for rec in dataset:
print(rec.subject, rec.raw.info['sfreq'])
If you use this dataset in your research, please cite the original authors.
BibTeX
@dataset{nm000402,
title = {Multi-Expert Seizure Annotation},
author = {William K. Ojemann and Daniel Zhou and Caren Armstrong and Nishant Sinha and Brian Litt and Erin Conrad},
doi = {10.82901/nemar.nm000402},
url = {https://doi.org/10.82901/nemar.nm000402},
}
About This Dataset#
Automated seizure detection and localization from intracranial EEG requires validated benchmark datasets with expert annotations, yet existing open datasets lack multi-expert consensus annotations and exclude stimulation-induced seizures. We present stereotactic EEG recordings from 78 seizures (44 spontaneous, 34 stimulation-induced) across 29 patients (19 from the University of Pennsylvania, 10 from the Children’s Hospital of Philadelphia) with drug-resistant epilepsy. Three or five board-certified epileptologists independently annotated each seizure for onset time, onset channels, and channels seizing at 10 seconds post-onset using a standardized protocol. All data follow Brain Imaging Data Structure (BIDS) standards and include electrode localizations, patient demographics, and clinical outcomes. This dataset enables the validation of seizure onset and spread detection and localization against human expert performance and supports comparative analysis of seizure networks across spontaneous and stimulation-induced seizures.
This NEMAR dataset redistributes Pennsieve Discover dataset 530, version 3 (doi:10.26275/dely-fzwx, published 2026-07-29 by
Penn CNT; licence on the record: “Creative Commons Attribution”; licence in the authors’ dataset_description.json: “CC-BY 4.0”). The release was already organised in BIDS by the authors (MNE-BIDS 0.17.0). Contents: 29 participants, 78 SEEG seizure recordings (EDF), 6.8 h in total, sampled at 512-2048 Hz.
Multi-Expert Seizure Annotation (spontaneous and stimulation-induced seizures in SEEG)
Overview
Contents
sub-<HUP|CHOP>###/ses-postimplant/ieeg/: one EDF per seizure, with_channels.tsv,_events.tsv/jsonand_ieeg.json.
View full README
Multi-Expert Seizure Annotation (spontaneous and stimulation-induced seizures in SEEG)
Overview
Contents
sub-<HUP|CHOP>###/ses-postimplant/ieeg/: one EDF per seizure, with_channels.tsv,_events.tsv/jsonand_ieeg.json. The task label encodes the seizure type and the approximate onset time in seconds from the start of the implant recording, as named by the authors (task-ictal<onset>: spontaneous seizure;task-stim<onset>: stimulation-induced seizure).participants.tsv/json: age, sex, mesial-temporal epilepsy, unifocal, lesional, Engel outcome, follow-up, disease duration, age at onset, number of seizures and number of stimulation-induced seizures (definitions in participants.json).annotations.tsv/json: the multi-expert annotations, one row per seizure: per-clinician unequivocal electrographic onset (UEO) time, UEO channels, channels seizing 10 s after onset, the consensus values, stimulation parameters, semiology, LVFA at onset (column definitions in annotations.json). Listed in.bidsignorebecause BIDS has no slot for dataset-level annotation tables.sub-*/ses-postimplant/ieeg/*_space-Other_electrodes.tsv/.json+_coordsystem.json: BIDS electrode files generated from the authors’ tables (HUP: native-space mm coordinates; CHOP: voxel indices only, units n/a). Columnlabelrenamedname,size= n/a, other columns kept.derivatives/electrode-localization/: the authors’ per-subject electrode localisation tables (HUP: native mm, tkrRAS and voxel coordinates with DK ROI; CHOP: voxel coordinates, tissue class, DK ROI), moved byte-identically from the release’s non-BIDSsub-*/derivatives/folders (mapping insourcedata/pennsieve-530-v3/moved_files.tsv).sourcedata/pennsieve-530-v3/: the Pennsieve record files (readme.md, manifest.json, changelog.md, banner.jpg), the release’s originalREADMEanddataset_description.json, the Pennsieve file listing with checksums, and our download verification (every file matched the Pennsieve SHA-256).
Changes made for NEMAR
No signal file (EDF), channels, events or ieeg sidecar was modified. Changes:
- the authors’ electrode tables were moved byte-identically to derivatives/electrode-localization/ (see above), and BIDS
_space-Other_electrodes.tsv/.json+_coordsystem.jsonwere generated from them (HUP: native-space mm coordinates; CHOP: voxel indices only, units n/a;labelrenamedname,size= n/a; the authors’matterLevels map moved into the column Description because the values use other spellings such as ‘grey’);
participants.jsonoutcome: the integer Levels map was moved into the Description (values are Engel subclasses such as 1.1, which the validator rejects against integer levels); participants.tsv values unchanged;dataset_description.json: License written as SPDXCC-BY-4.0, Authors taken from the Pennsieve record’s contributor list, Pennsieve DOI added to ReferencesAndLinks and HowToAcknowledge;this README replaces the release README (kept verbatim below and in sourcedata).
Privacy
EDF headers carry no patient information (patient field X X X X, start date 01.01.85 set by the authors); scans.tsv
acq_time is n/a. Subject labels are the authors’ study codes. Clinician names are replaced by Clin # in the release.
Ethics approval
From the authors’ dataset_description.json (EthicsApprovals): “University of Pennsylvania Human Research Protections Program, Institutional Review Boards (Protocol 703979, 811097, and/or 821778)”.
Funding
As listed by the authors in dataset_description.json (Funding).
How to cite
Cite the dataset (doi:10.26275/dely-fzwx) and the manuscript https://doi.org/10.64898/2026.01.15.26344025.
Licence
Creative Commons Attribution 4.0 (CC-BY-4.0), as released by the authors.
Original release README
References
Appelhoff, S., Sanderson, M., Brooks, T., Vliet, M., Quentin, R., Holdgraf, C., Chaumon, M., Mikulan, E., Tavabi, K., Höchenberger, R., Welke, D., Brunner, C., Rockhill, A., Larson, E., Gramfort, A. and Jas, M. (2019). MNE-BIDS: Organizing electrophysiological data into the BIDS format and facilitating their analysis. Journal of Open Source Software 4: (1896).https://doi.org/10.21105/joss.01896 Holdgraf, C., Appelhoff, S., Bickel, S., Bouchard, K., D’Ambrosio, S., David, O., … Hermes, D. (2019). iEEG-BIDS, extending the Brain Imaging Data Structure specification to human intracranial electrophysiology. Scientific Data, 6, 102. https://doi.org/10.1038/s41597-019-0105-7
Additional metadata and localisation (added 2026-10-08)
Compiled after the upload from the article, its supplement and the source deposit (each statement names its source). Text and sidecar metadata only; no data file was changed. Stimulation. HUP: bipolar, biphasic stimulation at 1 Hz, 3 mA, 300 µs (first 14 patients) or 500 µs pulse width; CHOP: 1–8 mA, 1–2 Hz, 300–500 µs (doi:10.64898/2026.01.15.26344025 (medRxiv, Ojemann et al. 2026), Methods). Per-seizure stim channels/parameters are in annotations.tsv (stim_channels, stim_frequency, stim_amplitude, stim_pulse_width). Recording system.**Natus Quantum system, sampling rate 512–2048 Hz (doi:10.64898/2026.01.15.26344025 (medRxiv, Ojemann et al. 2026), Methods); the deposited EDFs are at 1024 Hz with channels.tsv low_cutoff 0.0 / high_cutoff 512.0 (deposit *_ieeg.json SamplingFrequency, *_channels.tsv). PowerLineFrequency, SoftwareFilters and Manufacturer are n/a in the sidecars. **Reference scheme. Referenced to a contact hypothesised to be in non-epileptogenic tissue, typically medullary bone (doi:10.64898/2026.01.15.26344025 (medRxiv, Ojemann et al. 2026), Methods); iEEGReference is n/a in the deposit sidecars. Electrode types. sEEG depth electrodes; HUP: Ad-Tech (Oak Creek, WI); CHOP: PMT (MN, USA) or DIXI Medical (doi:10.64898/2026.01.15.26344025 (medRxiv, Ojemann et al. 2026), Methods). Localisation method.**CHOP: post-implant CT co-registered to pre-implant MRI with Gardel, Desikan–Killiany (DK) atlas segmentation in FreeSurfer. HUP: iEEG-recon, contacts localised in pre-implant T1 space and MNI152 space, FreeSurfer DK parcellation (doi:10.64898/2026.01.15.26344025 (medRxiv, Ojemann et al. 2026), Methods “Electrode Localization”). The deposit gives per-subject ``derivatives/electrodes.tsv``: CHOP files have voxel x/y/z + matter + DK brain_area; HUP files have mm_x/y/z, surfmm_x/y/z, vox_x/y/z, roi (FreeSurfer aseg/DK name) and roiNum. The MNI coordinates mentioned in the paper are not** in the deposit; the mm coordinates are subject-native (deposit does not name the space).
NEMAR Metadata#
[](https://doi.org/10.82901/nemar.nm000402) # Multi-Expert Seizure Annotation (spontaneous and stimulation-induced seizures in SEEG) ## Overview Automated seizure detection and localization from intracranial EEG requires validated benchmark datasets with expert annotations, yet existing open datasets lack multi-expert consensus annotations and exclude stimulation-induced seizures. We present stereotactic EEG recordings from 78 seizures (44 spontaneous, 34 stimulation-induced) across 29 patients (19 from the University of Pennsylvania, 10 from the Children’s Hospital of Philadelphia) with drug-resistant epilepsy. Three or five board-certified epileptologists independently annotated each seizure for onset time, onset channels, and channels seizing at 10 seconds post-onset using a standardized protocol. All data follow Brain Imaging Data Structure (BIDS) standards and include electrode localizations, patient demographics, and clinical outcomes. This dataset enables the validation of seizure onset and spread detection and localization against human expert performance and supports comparative analysis of seizure networks across spontaneous and stimulation-induced seizures. This NEMAR dataset redistributes Pennsieve Discover dataset 530, version 3 (doi:10.26275/dely-fzwx, published 2026-07-29 by Penn CNT; licence on the record: “Creative Commons Attribution”; licence in the authors’ dataset_description.json: “CC-BY 4.0”). The release was already organised in BIDS by the authors (MNE-BIDS 0.17.0). Contents: 29 participants, 78 SEEG seizure recordings (EDF), 6.8 h in total, sampled at 512-2048 Hz. ## Contents - sub-<HUP|CHOP>###/ses-postimplant/ieeg/: one EDF per seizure, with _channels.tsv, _events.tsv/json and _ieeg.json.
The task label encodes the seizure type and the approximate onset time in seconds from the start of the implant recording, as named by the authors (task-ictal<onset>: spontaneous seizure; task-stim<onset>: stimulation-induced seizure).
participants.tsv/json: age, sex, mesial-temporal epilepsy, unifocal, lesional, Engel outcome, follow-up, disease duration, age at onset, number of seizures and number of stimulation-induced seizures (definitions in participants.json).
annotations.tsv/json: the multi-expert annotations, one row per seizure: per-clinician unequivocal electrographic onset (UEO) time, UEO channels, channels seizing 10 s after onset, the consensus values, stimulation parameters, semiology, LVFA at onset (column definitions in annotations.json). Listed in .bidsignore because BIDS has no slot for dataset-level annotation tables.
sub-*/ses-postimplant/ieeg/*_space-Other_electrodes.tsv/.json + _coordsystem.json: BIDS electrode files generated from the authors’ tables (HUP: native-space mm coordinates; CHOP: voxel indices only, units n/a). Column label renamed name, size = n/a, other columns kept.
derivatives/electrode-localization/: the authors’ per-subject electrode localisation tables (HUP: native mm, tkrRAS and voxel coordinates with DK ROI; CHOP: voxel coordinates, tissue class, DK ROI), moved byte-identically from the release’s non-BIDS sub-*/derivatives/ folders (mapping in sourcedata/pennsieve-530-v3/moved_files.tsv).
sourcedata/pennsieve-530-v3/: the Pennsieve record files (readme.md, manifest.json, changelog.md, banner.jpg), the release’s original README and dataset_description.json, the Pennsieve file listing with checksums, and our download verification (every file matched the Pennsieve SHA-256).
## Changes made for NEMAR No signal file (EDF), channels, events or ieeg sidecar was modified. Changes: - the authors’ electrode tables were moved byte-identically to derivatives/electrode-localization/ (see above), and BIDS
_space-Other_electrodes.tsv/.json + _coordsystem.json were generated from them (HUP: native-space mm coordinates; CHOP: voxel indices only, units n/a; label renamed name, size = n/a; the authors’ matter Levels map moved into the column Description because the values use other spellings such as ‘grey’);
participants.json outcome: the integer Levels map was moved into the Description (values are Engel subclasses such as 1.1, which the validator rejects against integer levels); participants.tsv values unchanged;
dataset_description.json: License written as SPDX CC-BY-4.0, Authors taken from the Pennsieve record’s contributor list, Pennsieve DOI added to ReferencesAndLinks and HowToAcknowledge;
this README replaces the release README (kept verbatim below and in sourcedata).
## Privacy EDF headers carry no patient information (patient field X X X X, start date 01.01.85 set by the authors); scans.tsv acq_time is n/a. Subject labels are the authors’ study codes. Clinician names are replaced by Clin # in the release. ## Ethics approval From the authors’ dataset_description.json (EthicsApprovals): “University of Pennsylvania Human Research Protections Program, Institutional Review Boards (Protocol 703979, 811097, and/or 821778)”. ## Funding As listed by the authors in dataset_description.json (Funding). ## How to cite Cite the dataset (doi:10.26275/dely-fzwx) and the manuscript https://doi.org/10.64898/2026.01.15.26344025. ## Licence Creative Commons Attribution 4.0 (CC-BY-4.0), as released by the authors. ## Original release README References ———- Appelhoff, S., Sanderson, M., Brooks, T., Vliet, M., Quentin, R., Holdgraf, C., Chaumon, M., Mikulan, E., Tavabi, K., Höchenberger, R., Welke, D., Brunner, C., Rockhill, A., Larson, E., Gramfort, A. and Jas, M. (2019). MNE-BIDS: Organizing electrophysiological data into the BIDS format and facilitating their analysis. Journal of Open Source Software 4: (1896).https://doi.org/10.21105/joss.01896 Holdgraf, C., Appelhoff, S., Bickel, S., Bouchard, K., D’Ambrosio, S., David, O., … Hermes, D. (2019). iEEG-BIDS, extending the Brain Imaging Data Structure specification to human intracranial electrophysiology. Scientific Data, 6, 102. https://doi.org/10.1038/s41597-019-0105-7 ## Additional metadata and localisation (added 2026-10-08) Compiled after the upload from the article, its supplement and the source deposit (each statement names its source). Text and sidecar metadata only; no data file was changed. Stimulation. HUP: bipolar, biphasic stimulation at 1 Hz, 3 mA, 300 µs (first 14 patients) or 500 µs pulse width; CHOP: 1–8 mA, 1–2 Hz, 300–500 µs (doi:10.64898/2026.01.15.26344025 (medRxiv, Ojemann et al. 2026), Methods). Per-seizure stim channels/parameters are in annotations.tsv (stim_channels, stim_frequency, stim_amplitude, stim_pulse_width). Recording system. Natus Quantum system, sampling rate 512–2048 Hz (doi:10.64898/2026.01.15.26344025 (medRxiv, Ojemann et al. 2026), Methods); the deposited EDFs are at 1024 Hz with channels.tsv low_cutoff 0.0 / high_cutoff 512.0 (deposit _ieeg.json SamplingFrequency, *_channels.tsv). PowerLineFrequency, SoftwareFilters and Manufacturer are n/a in the sidecars. **Reference scheme.* Referenced to a contact hypothesised to be in non-epileptogenic tissue, typically medullary bone (doi:10.64898/2026.01.15.26344025 (medRxiv, Ojemann et al. 2026), Methods); iEEGReference is n/a in the deposit sidecars. Electrode types. sEEG depth electrodes; HUP: Ad-Tech (Oak Creek, WI); CHOP: PMT (MN, USA) or DIXI Medical (doi:10.64898/2026.01.15.26344025 (medRxiv, Ojemann et al. 2026), Methods). Localisation method. CHOP: post-implant CT co-registered to pre-implant MRI with Gardel, Desikan–Killiany (DK) atlas segmentation in FreeSurfer. HUP: iEEG-recon, contacts localised in pre-implant T1 space and MNI152 space, FreeSurfer DK parcellation (doi:10.64898/2026.01.15.26344025 (medRxiv, Ojemann et al. 2026), Methods “Electrode Localization”). The deposit gives per-subject derivatives/electrodes.tsv: CHOP files have voxel x/y/z + matter + DK brain_area; HUP files have mm_x/y/z, surfmm_x/y/z, vox_x/y/z, roi (FreeSurfer aseg/DK name) and roiNum. The MNI coordinates mentioned in the paper are not in the deposit; the mm coordinates are subject-native (deposit does not name the space).
License: CC-BY-4.0
Authors:
William K. Ojemann
Daniel Zhou
Caren Armstrong
Nishant Sinha
Brian Litt
… and 1 more
Versions:
Version |
DOI |
Released |
|---|---|---|
|
Cohort#
Dataset Statistics#
Age distribution by gender (n=28, range 4–58 yr, mean 29.8 yr)
Sex composition
Channel counts (ch)
Sampling frequencies (Hz)
Total recording duration: 6 h 47 min
Signal · Electrodes & live trace#
Live trace viewer — sub-CHOP005 · ses-postimplant · task-stim68881 · run-00
Showing one representative recording out of
29 subjects and 78 recordings in this dataset.
Browse the full set on OpenNeuro;
drop any other _ieeg.{set,edf,bdf,vhdr} file onto the
viewer (or pass ?ieeg=<url>) to inspect it.
Electrode layout — iEEG · 60 sensors — 60 channels
NEMAR Processing Statistics#
The plots below are generated by NEMAR’s automated EEG pipeline. The histogram shows pipeline success for data cleaning and ICA decomposition, the percentage of data frames and EEG channels retained after artefact removal, line noise per channel (RMS, dB), and the age/gender distribution of participants.
HED event descriptors word cloud
Manifest#
File Explorer#
Browse the BIDS file structure of this dataset. Records are fetched on demand from the EEGDash catalog the first time you open the explorer.
Full dataset metadata table
Dataset ID |
|
Title |
Multi-Expert Seizure Annotation |
Author (year) |
— |
Canonical |
— |
Importable as |
|
Year |
2019 |
Authors |
William K. Ojemann, Daniel Zhou, Caren Armstrong, Nishant Sinha, Brian Litt, Erin Conrad |
License |
CC-BY-4.0 |
Citation / DOI |
|
Source links |
OpenNeuro | NeMAR | Source URL |
Copy-paste BibTeX
@dataset{nm000402,
title = {Multi-Expert Seizure Annotation},
author = {William K. Ojemann and Daniel Zhou and Caren Armstrong and Nishant Sinha and Brian Litt and Erin Conrad},
doi = {10.82901/nemar.nm000402},
url = {https://doi.org/10.82901/nemar.nm000402},
}
API Reference#
eegdash.datasetEEGDashDataset- class eegdash.dataset.NM000402(cache_dir: str, query: dict | None = None, s3_bucket: str | None = None, **kwargs)[source]#
Multi-Expert Seizure Annotation
- Study:
nm000402(NeMAR)- Author (year):
—
- Canonical:
—
Also importable as:
NM000402.Modality:
ieeg; Subject type:Unknown. Subjects: 29; recordings: 78; tasks: 78.- Parameters:
cache_dir (str | Path) – Directory where data are cached locally.
query (dict | None) – Additional MongoDB-style filters to AND with the dataset selection. Must not contain the key
dataset.s3_bucket (str | None) – Base S3 bucket used to locate the data.
**kwargs (dict) – Additional keyword arguments forwarded to
EEGDashDataset.
- data_dir#
Local dataset cache directory (
cache_dir / dataset_id).- Type:
Path
Notes
Each item is a recording; recording-level metadata are available via
dataset.description.querysupports MongoDB-style filters on fields inALLOWED_QUERY_FIELDSand is combined with the dataset filter. Dataset-specific caveats are not provided in the summary metadata.References
OpenNeuro dataset: https://openneuro.org/datasets/nm000402 NeMAR dataset: https://nemar.org/dataexplorer/detail?dataset_id=nm000402 DOI: https://doi.org/10.82901/nemar.nm000402
Examples
>>> from eegdash.dataset import NM000402 >>> dataset = NM000402(cache_dir="./data") >>> recording = dataset[0] >>> raw = recording.load()
- __init__(cache_dir: str, query: dict | None = None, s3_bucket: str | None = None, **kwargs)[source]#
- save(path: str, overwrite: bool = False, offset: int = 0)[source]#
Save datasets to files by creating one subdirectory for each dataset:
path/ 0/ 0-raw.fif | 0-epo.fif description.json raw_preproc_kwargs.json (if raws were preprocessed) window_kwargs.json (if this is a windowed dataset) window_preproc_kwargs.json (if windows were preprocessed) target_name.json (if target_name is not None and dataset is raw) 1/ 1-raw.fif | 1-epo.fif description.json raw_preproc_kwargs.json (if raws were preprocessed) window_kwargs.json (if this is a windowed dataset) window_preproc_kwargs.json (if windows were preprocessed) target_name.json (if target_name is not None and dataset is raw)
- Parameters:
path (str) –
- Directory in which subdirectories are created to store
-raw.fif | -epo.fif and .json files to.
overwrite (bool) – Whether to delete old subdirectories that will be saved to in this call.
offset (int) – If provided, the integer is added to the id of the dataset in the concat. This is useful in the setting of very large datasets, where one dataset has to be processed and saved at a time to account for its original position.
BaseDataset from braindecode — windowed via create_windows_from_events.braindecodeDataLoader; supports parallel workers and on-the-fly augmentations.pytorchSwap any load_dataset(...) call for nm000402 to reproduce the tutorial on this dataset.
Citation
William K. Ojemann, Daniel Zhou, Caren Armstrong, Nishant Sinha, Brian Litt, … (2019). Multi-Expert Seizure Annotation. 10.82901/nemar.nm000402
Provenance
¹Contributed to nemar in BIDS format.
²Curated & ingested by the EEGDash catalog; see CITATION.cff for canonical reference.
³Persistent identifier: 10.82901/nemar.nm000402.
See Also#
eegdash.dataset.EEGDashDataseteegdash.dataset