EEGdash›NeMAR›NM000410
Iss. 410 · 79 subjects · 79 recordings · CC-BY-NC-SA-4.0
Dataset Brief · Vanderbilt BIEN Lab SOZ Resting State SEEG Classification (pr…

NM000410: ieeg dataset, 79 subjects#

Vanderbilt BIEN Lab SOZ Resting State SEEG Classification (processed resting-state SEEG)

Access recordings and metadata through EEGDash.

Citation: Kate Wang, Sameer Sundrani, Dario Englot (2025). Vanderbilt BIEN Lab SOZ Resting State SEEG Classification (processed resting-state SEEG). 10.82901/nemar.nm000410

Modality: ieeg Subjects: 79 Recordings: 79 License: CC-BY-NC-SA-4.0 Source: nemar

Metadata: Complete (100%)

79-participant iEEG dataset — Vanderbilt BIEN Lab SOZ Resting State SEEG Classification (processed resting-state SEEG).

iEEG · 58 (3), 90 (3), 42 (3), 69 (2), 70 (2), 109 (2), 103 (2), 111 (2), 124 (2), 95 (2), 73 (2), 52 (2), 76 (2), 59 (2), 94 (2), 100 (2), 113 (2), 126 (2), 87, 46, 75, 141, 96, 43, 98, 86, 110, 53, 118, 71, 51, 134, 65, 92, 67, 77, 63, 83, 112, 105, 72, 120, 117, 26, 136, 89, 106, 108, 93, 60, 85, 84, 68, 38, 66, 55, 74, 133 ch500, 512, 1024, 2048 HzBIDS 1.10.0Task · rest
Layer 01Study
What was asked
Hypothesis, independent & dependent variables, paradigm, cohort, and the editorial caveats around what the recordings can and cannot answer.
Layer 02Signal · BIDS
What was recorded
Sidecars, channels & electrodes, coordinate system, event semantics, and quality stats from the NEMAR pipeline when available.
Layer 03Training · ML
What you can train on
Recommended access modes — MNE Raw, braindecode windows, PyTorch DataLoader — plus the targets the metadata makes addressable.
§ 01Access · Get started

Quickstart#

Install

pip install eegdash

Access the data

from eegdash.dataset import NM000410

dataset = NM000410(cache_dir="./data")
# Get the raw object of the first recording
raw = dataset.datasets[0].raw
print(raw.info)

Filter by subject

dataset = NM000410(cache_dir="./data", subject="01")

Advanced query

dataset = NM000410(
    cache_dir="./data",
    query={"subject": {"$in": ["01", "02"]}},
)

Iterate recordings

for rec in dataset:
    print(rec.subject, rec.raw.info['sfreq'])

If you use this dataset in your research, please cite the original authors.

BibTeX

@dataset{nm000410,
  title = {Vanderbilt BIEN Lab SOZ Resting State SEEG Classification (processed resting-state SEEG)},
  author = {Kate Wang and Sameer Sundrani and Dario Englot},
  doi = {10.82901/nemar.nm000410},
  url = {https://doi.org/10.82901/nemar.nm000410},
}
§ 02Study · The README

About This Dataset#

NEMAR redistribution of Pennsieve Discover dataset 492, version 1 (doi:10.26275/wtaz-dbst; “Vanderbilt BIEN Lab SOZ Resting

State SEEG Classification”; licence on the record: “Creative Commons Attribution - NonCommercial-ShareAlike”, written here as CC-BY-NC-SA-4.0). It accompanies Sundrani et al. (2025), “Deep learning on brief interictal intracranial recordings can accurately characterize seizure onset zones”, Epilepsia 66(9):3180-3192, doi:10.1111/epi.18478. This is a processed-data (derivative) dataset: the signals are the authors’ filtered bipolar SEEG; raw recordings are not part of the release.

Release description (verbatim): This repository contains a de-identified codebase for training and analyzing deep neural networks that classify seizure onset zones (SOZ) vs. non-SOZ brain regions using resting-state stereoelectroencephalography (SEEG) recordings.

The citation for the associated manuscript is listed below:

DOI

Vanderbilt resting-state SEEG for seizure-onset-zone classification (derivative dataset)

Overview

Sundrani, S., Johnson, G. W., Doss, D. J., Makhoul, G. S., Hidalgo Monroy Lerma, B., Reda, A., Cavender, A. C., Liao, E., Rogers, B. P., Williams Roberson, S., Bick, S. K., Morgan, V. L., & Englot, D. J. (2025). Deep learning on brief interictal intracranial recordings can accurately characterize seizure onset zones. Epilepsia, 66(9), 3180–3192. Portico. https://doi.org/10.1111/epi.18478

Contents

View full README

DOI

Vanderbilt resting-state SEEG for seizure-onset-zone classification (derivative dataset)

Overview

Sundrani, S., Johnson, G. W., Doss, D. J., Makhoul, G. S., Hidalgo Monroy Lerma, B., Reda, A., Cavender, A. C., Liao, E., Rogers, B. P., Williams Roberson, S., Bick, S. K., Morgan, V. L., & Englot, D. J. (2025). Deep learning on brief interictal intracranial recordings can accurately characterize seizure onset zones. Epilepsia, 66(9), 3180–3192. Portico. https://doi.org/10.1111/epi.18478

Contents

  • 79 participants, one 5-minute recording each (sub-<code>/ieeg/sub-<code>_task-rest_ieeg.vhdr), BrainVision, float32, microvolts, 500-2048 Hz as released. Participant codes are the release codes (Epat##, Spat##, pat##; the release does not explain the prefixes). The paper reports 78 patients; the release has 79 files. Codes not in any cross-validation test fold of the authors’ results: [‘Epat09’].

  • _channels.tsv: bipolar channel, region (release bip_montage_region), and soz_region_label: the region-level SOZ ground truth used by the authors (‘true_label’ in the results JSON). See channels.json.

  • _electrodes.tsv + _coordsystem.json: one row per bipolar channel with its region; the release gives no coordinates (x/y/z n/a).

  • participants.tsv: release code and the authors’ test fold.

  • Unreadable samples in the release: Epat09: rows 46875-46999 (93.750-94.000 s), channels 0-64 (one corrupt gzip chunk in the HDF5 file; the file matches the Pennsieve SHA-256, so the release itself carries it; the samples are NaN here and marked in that recording’s _events.tsv)

  • sourcedata/pennsieve-492-v1/: the complete release byte-identical (MAT files, code, model weights, results JSON, location_accuracy.csv, README.md, Pennsieve readme/manifest/changelog/banner), except macOS .DS_Store files and a compiled .pyc (1 files skipped: [‘files/models/__pycache__/multi_scale_ori.cpython-39.pyc’]); the Pennsieve listing with checksums; roundtrip_float32.tsv (max relative float32 error over all files: 5.77e-08).

Recording and processing (from the paper)

Five minutes of resting-state (interictal, non-SPES) SEEG from patients with drug-resistant epilepsy evaluated at Vanderbilt University Medical Center; patients awake, eyes closed, told to try not to fall asleep; recordings at least 4 h apart from electroclinical activity. Filtering: MATLAB filtfilt Butterworth passbands 1-59, 61-119 and 121-150 Hz.

Contacts were assigned to Desikan-Killiany regions; SOZs were defined on a region basis as regions containing any contact involved in the ictal onset of one or more seizures. Cohort (paper Table 1, n = 78): 45 female, 33 male; per-patient age and sex are not in the release.

Changes made for NEMAR

filt_data (float64, volts) written as float32 microvolts (BrainVision); channel names without spaces; unreadable source samples written as NaN (see Contents). Nothing else changed.

Privacy

The release contains only signals, channel/region labels, model outputs and code. No names, dates or hospital numbers were found in the MAT strings, JSON or text files. The code README lists the corresponding author’s institutional e-mail.

Ethics approval

Sundrani et al. 2025, Methods (Participants and Resting State SEEG): “This study was approved by the Vanderbilt Institutional Review Board, and informed subject consent was obtained.”

Funding

Sundrani et al. 2025, Acknowledgements (verbatim in dataset_description.json).

Licence

CC-BY-NC-SA-4.0 (Pennsieve licence field: “Creative Commons Attribution - NonCommercial-ShareAlike”). Non-commercial use only.

Additional metadata and localisation (added 2026-10-08)

Compiled after the upload from the article, its supplement and the source deposit (each statement names its source). Text and sidecar metadata only; no data file was changed. Recording system. Amplifier not stated (n/a). sampling_freq in the .mat files: 512 Hz (40), 500 Hz (20), 1024 Hz (14), 2048 Hz (5); filt_data holds 150000-153600 samples per channel (deposit, Voyager Job). The deposit README states inputs at 500-512 Hz (data are resampled to 500 Hz by the code, resample_to_500hz). Filtering: MATLAB filtfilt Butterworth passbands 1-59, 61-119 and 121-150 Hz (paper Methods “Data Preprocessing”) - the stored filt_data are the filtered signals (variable name; deposit code). Reference scheme. Bipolar montage of adjacent contacts (bip_montage_label e.g. “LAC1 - LAC2”; deposit .mat). Original recording reference not stated. Electrode types. SEEG depth electrodes; manufacturer not stated. Localisation method. Contacts localised on post-implantation CT with CRAnial Vault Explorer (CRAVE) and each contact assigned to a Desikan-Killiany (DK) region; verified by a staff engineer, attending neurosurgeon and attending epileptologist. SOZ defined per DK region containing any contact involved in ictal onset of >=1 seizure (paper Methods). Coordinates are not in the deposit; only DK labels per bipolar channel. Cohort (paper Table 1). n=78; female 45 (57.7%); age mean 34.6 (SD 12.4); outcome at one year: Engel I 23, II 5, III 7, IV 3, neuromodulation responder 15, non-responder 10, none 15.

Each electrodes.tsv now has a soz_region_label column (yes/no/n/a): whether the channel’s DK region is one of the patient’s SOZ regions in the deposit results JSON.

NEMAR Metadata#

[![DOI](https://img.shields.io/badge/DOI-10.82901%2Fnemar.nm000410-blue)](https://doi.org/10.82901/nemar.nm000410) # Vanderbilt resting-state SEEG for seizure-onset-zone classification (derivative dataset) ## Overview NEMAR redistribution of Pennsieve Discover dataset 492, version 1 (doi:10.26275/wtaz-dbst; “Vanderbilt BIEN Lab SOZ Resting State SEEG Classification”; licence on the record: “Creative Commons Attribution - NonCommercial-ShareAlike”, written here as CC-BY-NC-SA-4.0). It accompanies Sundrani et al. (2025), “Deep learning on brief interictal intracranial recordings can accurately characterize seizure onset zones”, Epilepsia 66(9):3180-3192, doi:10.1111/epi.18478. This is a processed-data (derivative) dataset: the signals are the authors’ filtered bipolar SEEG; raw recordings are not part of the release. Release description (verbatim): This repository contains a de-identified codebase for training and analyzing deep neural networks that classify seizure onset zones (SOZ) vs. non-SOZ brain regions using resting-state stereoelectroencephalography (SEEG) recordings. The citation for the associated manuscript is listed below: Sundrani, S., Johnson, G. W., Doss, D. J., Makhoul, G. S., Hidalgo Monroy Lerma, B., Reda, A., Cavender, A. C., Liao, E., Rogers, B. P., Williams Roberson, S., Bick, S. K., Morgan, V. L., & Englot, D. J. (2025). Deep learning on brief interictal intracranial recordings can accurately characterize seizure onset zones. Epilepsia, 66(9), 3180–3192. Portico. https://doi.org/10.1111/epi.18478 ## Contents - 79 participants, one 5-minute recording each (sub-<code>/ieeg/sub-<code>_task-rest_ieeg.vhdr), BrainVision,

float32, microvolts, 500-2048 Hz as released. Participant codes are the release codes (Epat##, Spat##, pat##; the release does not explain the prefixes). The paper reports 78 patients; the release has 79 files. Codes not in any cross-validation test fold of the authors’ results: [‘Epat09’].

  • _channels.tsv: bipolar channel, region (release bip_montage_region), and soz_region_label: the region-level SOZ ground truth used by the authors (‘true_label’ in the results JSON). See channels.json.

  • _electrodes.tsv + _coordsystem.json: one row per bipolar channel with its region; the release gives no coordinates (x/y/z n/a).

  • participants.tsv: release code and the authors’ test fold.

  • Unreadable samples in the release: Epat09: rows 46875-46999 (93.750-94.000 s), channels 0-64 (one corrupt gzip chunk in the HDF5 file; the file matches the Pennsieve SHA-256, so the release itself carries it; the samples are NaN here and marked in that recording’s _events.tsv)

  • sourcedata/pennsieve-492-v1/: the complete release byte-identical (MAT files, code, model weights, results JSON, location_accuracy.csv, README.md, Pennsieve readme/manifest/changelog/banner), except macOS .DS_Store files and a compiled .pyc (1 files skipped: [‘files/models/__pycache__/multi_scale_ori.cpython-39.pyc’]); the Pennsieve listing with checksums; roundtrip_float32.tsv (max relative float32 error over all files: 5.77e-08).

## Recording and processing (from the paper) Five minutes of resting-state (interictal, non-SPES) SEEG from patients with drug-resistant epilepsy evaluated at Vanderbilt University Medical Center; patients awake, eyes closed, told to try not to fall asleep; recordings at least 4 h apart from electroclinical activity. Filtering: MATLAB filtfilt Butterworth passbands 1-59, 61-119 and 121-150 Hz. Contacts were assigned to Desikan-Killiany regions; SOZs were defined on a region basis as regions containing any contact involved in the ictal onset of one or more seizures. Cohort (paper Table 1, n = 78): 45 female, 33 male; per-patient age and sex are not in the release. ## Changes made for NEMAR filt_data (float64, volts) written as float32 microvolts (BrainVision); channel names without spaces; unreadable source samples written as NaN (see Contents). Nothing else changed. ## Privacy The release contains only signals, channel/region labels, model outputs and code. No names, dates or hospital numbers were found in the MAT strings, JSON or text files. The code README lists the corresponding author’s institutional e-mail. ## Ethics approval Sundrani et al. 2025, Methods (Participants and Resting State SEEG): “This study was approved by the Vanderbilt Institutional Review Board, and informed subject consent was obtained.” ## Funding Sundrani et al. 2025, Acknowledgements (verbatim in dataset_description.json). ## Licence CC-BY-NC-SA-4.0 (Pennsieve licence field: “Creative Commons Attribution - NonCommercial-ShareAlike”). Non-commercial use only. ## Additional metadata and localisation (added 2026-10-08) Compiled after the upload from the article, its supplement and the source deposit (each statement names its source). Text and sidecar metadata only; no data file was changed. Recording system. Amplifier not stated (n/a). sampling_freq in the .mat files: 512 Hz (40), 500 Hz (20), 1024 Hz (14), 2048 Hz (5); filt_data holds 150000-153600 samples per channel (deposit, Voyager Job). The deposit README states inputs at 500-512 Hz (data are resampled to 500 Hz by the code, resample_to_500hz). Filtering: MATLAB filtfilt Butterworth passbands 1-59, 61-119 and 121-150 Hz (paper Methods “Data Preprocessing”) - the stored filt_data are the filtered signals (variable name; deposit code). Reference scheme. Bipolar montage of adjacent contacts (bip_montage_label e.g. “LAC1 - LAC2”; deposit .mat). Original recording reference not stated. Electrode types. SEEG depth electrodes; manufacturer not stated. Localisation method. Contacts localised on post-implantation CT with CRAnial Vault Explorer (CRAVE) and each contact assigned to a Desikan-Killiany (DK) region; verified by a staff engineer, attending neurosurgeon and attending epileptologist. SOZ defined per DK region containing any contact involved in ictal onset of >=1 seizure (paper Methods). Coordinates are not in the deposit; only DK labels per bipolar channel. Cohort (paper Table 1). n=78; female 45 (57.7%); age mean 34.6 (SD 12.4); outcome at one year: Engel I 23, II 5, III 7, IV 3, neuromodulation responder 15, non-responder 10, none 15. Each electrodes.tsv now has a soz_region_label column (yes/no/n/a): whether the channel’s DK region is one of the patient’s SOZ regions in the deposit results JSON.

License: CC-BY-NC-SA-4.0

Authors:

  • Kate Wang

  • Sameer Sundrani

  • Dario Englot

Versions:

Version

DOI

Released

current

10.82901/nemar.nm000410

§ 03Cohort · Participants

Cohort#

Dataset Statistics#

Channel counts (ch)

263842434651525355585960636566676869707172737475767783848586878990929394959698100103105106108109110111112113117118120124126133134136141

Sampling frequencies (Hz)

50051210242048

Total recording duration: 6 h 35 min

§ 04Signal · Electrodes & trace

Signal · Electrodes & live trace#

Fig. 01 Signal & montage 58 (3), 90 (3), 42 (3), 69 (2), 70 (2), 109 (2), 103 (2), 111 (2), 124 (2), 95 (2), 73 (2), 52 (2), 76 (2), 59 (2), 94 (2), 100 (2), 113 (2), 126 (2), 87, 46, 75, 141, 96, 43, 98, 86, 110, 53, 118, 71, 51, 134, 65, 92, 67, 77, 63, 83, 112, 105, 72, 120, 117, 26, 136, 89, 106, 108, 93, 60, 85, 84, 68, 38, 66, 55, 74, 133 ch · iEEG · 500, 512, 1024, 2048 Hz · 79 subjects, 79 recordings
Live trace viewer — sub-Epat02 · task-rest

Showing one representative recording out of 79 subjects and 79 recordings in this dataset. Browse the full set on OpenNeuro; drop any other _ieeg.{set,edf,bdf,vhdr} file onto the viewer (or pass ?ieeg=<url>) to inspect it.

No scalp electrode layout is currently indexed for this dataset. Once the eegdash montage registry ingests it, the interactive viewer will appear here automatically.

NEMAR Processing Statistics#

The plots below are generated by NEMAR’s automated EEG pipeline. The histogram shows pipeline success for data cleaning and ICA decomposition, the percentage of data frames and EEG channels retained after artefact removal, line noise per channel (RMS, dB), and the age/gender distribution of participants.

HED event descriptors word cloud HED event descriptors word cloud — NM000410
§ 05Manifest · BIDS tree

Manifest#

File Explorer#

Browse the BIDS file structure of this dataset. Records are fetched on demand from the EEGDash catalog the first time you open the explorer.

Recordings—
Files—
Subjects—
Modalities—
Click to load file structure…
Full dataset metadata table

Dataset ID

NM000410

Title

Vanderbilt BIEN Lab SOZ Resting State SEEG Classification (processed resting-state SEEG)

Author (year)

—

Canonical

—

Importable as

NM000410

Year

2025

Authors

Kate Wang, Sameer Sundrani, Dario Englot

License

CC-BY-NC-SA-4.0

Citation / DOI

10.82901/nemar.nm000410

Source links

OpenNeuro | NeMAR | Source URL

Copy-paste BibTeX
@dataset{nm000410,
  title = {Vanderbilt BIEN Lab SOZ Resting State SEEG Classification (processed resting-state SEEG)},
  author = {Kate Wang and Sameer Sundrani and Dario Englot},
  doi = {10.82901/nemar.nm000410},
  url = {https://doi.org/10.82901/nemar.nm000410},
}
§ 06API · Programmatic access

API Reference#

Signature
eegdash.dataset
class
eegdash.dataset.NM000410(cache_dir, query=None, s3_bucket=None, **kwargs)
Bases: EEGDashDataset
Author (year)—
Canonical—
Importable asNM000410
Sourceeegdash/dataset/registry.py · [source ↗]
class eegdash.dataset.NM000410(cache_dir: str, query: dict | None = None, s3_bucket: str | None = None, **kwargs)[source]#

Vanderbilt BIEN Lab SOZ Resting State SEEG Classification (processed resting-state SEEG)

Study:

nm000410 (NeMAR)

Author (year):

—

Canonical:

—

Also importable as: NM000410.

Modality: ieeg; Subject type: Unknown. Subjects: 79; recordings: 79; tasks: 1.

Parameters:
  • cache_dir (str | Path) – Directory where data are cached locally.

  • query (dict | None) – Additional MongoDB-style filters to AND with the dataset selection. Must not contain the key dataset.

  • s3_bucket (str | None) – Base S3 bucket used to locate the data.

  • **kwargs (dict) – Additional keyword arguments forwarded to EEGDashDataset.

data_dir#

Local dataset cache directory (cache_dir / dataset_id).

Type:

Path

query#

Merged query with the dataset filter applied.

Type:

dict

records#

Metadata records used to build the dataset, if pre-fetched.

Type:

list[dict] | None

Notes

Each item is a recording; recording-level metadata are available via dataset.description. query supports MongoDB-style filters on fields in ALLOWED_QUERY_FIELDS and is combined with the dataset filter. Dataset-specific caveats are not provided in the summary metadata.

References

OpenNeuro dataset: https://openneuro.org/datasets/nm000410 NeMAR dataset: https://nemar.org/dataexplorer/detail?dataset_id=nm000410 DOI: https://doi.org/10.82901/nemar.nm000410

Examples

>>> from eegdash.dataset import NM000410
>>> dataset = NM000410(cache_dir="./data")
>>> recording = dataset[0]
>>> raw = recording.load()
__init__(cache_dir: str, query: dict | None = None, s3_bucket: str | None = None, **kwargs)[source]#
save(path: str, overwrite: bool = False, offset: int = 0)[source]#

Save datasets to files by creating one subdirectory for each dataset:

path/
    0/
        0-raw.fif | 0-epo.fif
        description.json
        raw_preproc_kwargs.json (if raws were preprocessed)
        window_kwargs.json (if this is a windowed dataset)
        window_preproc_kwargs.json  (if windows were preprocessed)
        target_name.json (if target_name is not None and dataset is raw)
    1/
        1-raw.fif | 1-epo.fif
        description.json
        raw_preproc_kwargs.json (if raws were preprocessed)
        window_kwargs.json (if this is a windowed dataset)
        window_preproc_kwargs.json  (if windows were preprocessed)
        target_name.json (if target_name is not None and dataset is raw)
Parameters:
  • path (str) –

    Directory in which subdirectories are created to store

    -raw.fif | -epo.fif and .json files to.

  • overwrite (bool) – Whether to delete old subdirectories that will be saved to in this call.

  • offset (int) – If provided, the integer is added to the id of the dataset in the concat. This is useful in the setting of very large datasets, where one dataset has to be processed and saved at a time to account for its original position.

Access modesMNE → braindecode → PyTorch → ML
.rawMNE Raw object — standard tools (filter, epoch, ICA, plot_psd).mne
DataLoaderWraps the windowed dataset into a PyTorch DataLoader; supports parallel workers and on-the-fly augmentations.pytorch
Zarr cacheOptional braindecode Zarr mirror for fast resume; persisted to cache_dir.zarr
Hugging FaceNo per-dataset mirror published yet — browse the EEGDash org listing for sibling datasets. See the datasets loader API.huggingface
Croissant 1.0Machine-readable JSON-LD descriptor — NM000410.croissant.json (MLCommons schema, ingestible by PyTorch / TensorFlow / JAX).mlcommons
Examples using EEGDashcurated · start here

Swap any load_dataset(...) call for nm000410 to reproduce the tutorial on this dataset.

Citation

Kate Wang, Sameer Sundrani, Dario Englot (2025). Vanderbilt BIEN Lab SOZ Resting State SEEG Classification (processed resting-state SEEG). 10.82901/nemar.nm000410

Provenance

¹Contributed to nemar in BIDS format.

²Curated & ingested by the EEGDash catalog; see CITATION.cff for canonical reference.

³Persistent identifier: 10.82901/nemar.nm000410.

BIDS
BIDS 1.10.0
Sidecars
channels · electrodes · coordsystem · eeg.json
Provenance
CC-BY-NC-SA-4.0 · 10.82901/nemar.nm000410
Machine-readable
Mirrors

See Also#