EEGdash›NeMAR›NM000365
Iss. 365 · 12 subjects · 67 recordings · CC-BY-4.0
Dataset Brief · Du-IN Chinese word-reading sEEG (Zheng et al., NeurIPS 2024)

NM000365: ieeg dataset, 12 subjects#

Du-IN Chinese word-reading sEEG (Zheng et al., NeurIPS 2024): authors’ preprocessed 3-s epochs, 12 subjects, 61 words

Access recordings and metadata through EEGDash.

Citation: Hui Zheng, Hai-Teng Wang, Wei-Bang Jiang, Zhong-Tao Chen, Li He, Pei-Yang Lin, Peng-Hu Wei, Guo-Guang Zhao, Yun-Zhe Liu (2021). Du-IN Chinese word-reading sEEG (Zheng et al., NeurIPS 2024): authors’ preprocessed 3-s epochs, 12 subjects, 61 words. 10.82901/nemar.nm000365

Modality: ieeg Subjects: 12 Recordings: 67 License: CC-BY-4.0 Source: nemar

Metadata: Complete (100%)

12-participant iEEG dataset — Du-IN Chinese word-reading sEEG (Zheng et al., NeurIPS 2024): authors' preprocessed 3-s epochs, 12 subjects, 61 words.

iEEG · 90 (6), 94 (6), 115 (6), 147 (6), 67 (6), 104 (6), 88 (6), 101 (6), 98 (6), 123 (5), 100 (5), 62 (3) ch1000 HzBIDS 1.10.0Task · wordreading
Layer 01Study
What was asked
Hypothesis, independent & dependent variables, paradigm, cohort, and the editorial caveats around what the recordings can and cannot answer.
Layer 02Signal · BIDS
What was recorded
Sidecars, channels & electrodes, coordinate system, event semantics, and quality stats from the NEMAR pipeline when available.
Layer 03Training · ML
What you can train on
Recommended access modes — MNE Raw, braindecode windows, PyTorch DataLoader — plus the targets the metadata makes addressable.
§ 01Access · Get started

Quickstart#

Install

pip install eegdash

Access the data

from eegdash.dataset import NM000365

dataset = NM000365(cache_dir="./data")
# Get the raw object of the first recording
raw = dataset.datasets[0].raw
print(raw.info)

Filter by subject

dataset = NM000365(cache_dir="./data", subject="01")

Advanced query

dataset = NM000365(
    cache_dir="./data",
    query={"subject": {"$in": ["01", "02"]}},
)

Iterate recordings

for rec in dataset:
    print(rec.subject, rec.raw.info['sfreq'])

If you use this dataset in your research, please cite the original authors.

BibTeX

@dataset{nm000365,
  title = {Du-IN Chinese word-reading sEEG (Zheng et al., NeurIPS 2024): authors' preprocessed 3-s epochs, 12 subjects, 61 words},
  author = {Hui Zheng and Hai-Teng Wang and Wei-Bang Jiang and Zhong-Tao Chen and Li He and Pei-Yang Lin and Peng-Hu Wei and Guo-Guang Zhao and Yun-Zhe Liu},
  doi = {10.82901/nemar.nm000365},
  url = {https://doi.org/10.82901/nemar.nm000365},
}
§ 02Study · The README

About This Dataset#

These are not raw recordings. This dataset is a BIDS packaging of the word-reading sEEG data released with

Du-IN: Discrete units-guided mask modeling for decoding speech from Intracranial Neural signals (Zheng, Wang, Jiang, Chen, He, Lin, Wei, Zhao, Liu; NeurIPS 2024; doi:10.52202/079017-2542; arXiv:2405.11459). The authors published only a preprocessed version of the 61-word reading data on Hugging Face (liulab-repository/Du-IN, revision 2c85b670230dced33f2f1cbd0903648680ecc7d0, last modified 2025-04-30, licence CC-BY-4.0). Their GitHub README says: “we can only share the preprocessed version of the 61-word reading sEEG dataset (~3 hours)”. The ~12 h of non-task recordings used for pre-training are not shared. dataset_description.json therefore declares DatasetType: derivative.

does not link age or sex to individual subjects, so participants.tsv holds only the subject IDs. 7 to 13 sEEG

electrodes per subject, placed for clinical reasons only. Recorded at 2000 Hz.

DOI

Du-IN Chinese word-reading sEEG (Zheng et al., NeurIPS 2024): the authors’ preprocessed epochs

  • Task: reading Chinese words aloud. The 61-word set and the authors’ English translations are in the paper (Table 4) and in the word_en column of the events files. Each trial starts with the word in white text (0.5 s). The text then turns green, which is the go cue, and stays on screen for 2 s while the subject reads the word aloud. A fixation cross follows for 0.5 s. Each ~10-minute block presents every word twice in random order (122 trials). The paper reports 25 blocks in total.

View full README

DOI

Du-IN Chinese word-reading sEEG (Zheng et al., NeurIPS 2024): the authors’ preprocessed epochs

  • Task: reading Chinese words aloud. The 61-word set and the authors’ English translations are in the paper (Table 4) and in the word_en column of the events files. Each trial starts with the word in white text (0.5 s). The text then turns green, which is the go cue, and stays on screen for 2 s while the subject reads the word aloud. A fixation cross follows for 0.5 s. Each ~10-minute block presents every word twice in random order (122 trials). The paper reports 25 blocks in total.

  • Ethics, as stated by the authors: “Experiments that contribute to this work were approved by IRB. All subjects consent to participate.” The consent form includes a privacy statement and data retention “after deleting all personal identification information (PII)”. The release contains no names, dates or imaging.

Additional study details (from the paper, arXiv:2405.11459v3: section 4.1, Appendix A, Ethics Statement, Appendix O)

  • Implant: 7 to 13 sEEG depth electrodes per subject; the paper gives 72 to 158 channels per subject at acquisition (the released bipolar channel counts are 62 to 147, see the table below). Electrodes carry 8-16 channels each (section 4.3). “All electrode locations are exclusively dictated by clinical considerations.” Most subjects had electrodes on one hemisphere only; the paper shows per-subject electrode locations as figures (Appendix O, Figures 9-11) but no coordinates were released.

  • Recordings: about 15 h of 2000 Hz sEEG per subject, of which about 3 h are task recordings (the part released here) and about 12 h are non-task recordings during wakefulness (not released). The subject’s voice was recorded simultaneously during the task; the audio is not part of the release.

  • Word set: 61 words chosen for (1) versatility in generating sentences, (2) expressing basic caregiving needs and (3) covering as many Chinese pronunciation combinations as possible. The design follows Moses et al. (2021).

  • Recording site: not stated explicitly by the paper or the release. Two authors (P.-H. Wei, G.-G. Zhao) are affiliated with Capital Medical University, Xuanwu Hospital, Beijing, and the release directory is named seeg.he2023xuanwu. No InstitutionName is set in the sidecars for this reason.

  • Amplifier/recording system, electrode manufacturer and acquisition reference are not stated in the paper or the release (left out of the sidecars).

  • Consent (Ethics Statement): adults with full civil capacity signed written informed consent; for minors or participants without full civil capacity, the legal guardian signed. The consent form covers research purpose, risks, data use, a privacy statement (PII not disclosed), data retention after deleting all PII, voluntariness and the right to withdraw at any time.

Preprocessing done by the authors (before release)

Paper section 4.2: band-pass 0.5-200 Hz, 50 Hz notch, resampling to 1000 Hz, bipolar re-reference. The signals were then cut into 3-s samples (3000 points), each labelled with its word. The released variant is named dataset.bipolar.default.unaligned. The paper also lists per-channel z-scoring, but the stored values are not z-scored: per-channel standard deviations are of order 1e-5. That step is presumably applied at training time.

The release does not state: - the unit (the magnitudes are only plausible as volts, so V is used and the numbers are left unchanged); - the contact pairing of each bipolar channel (channel names are as released); - where each 3-s epoch sits relative to word onset, go cue or speech onset; - whether the stored epoch order is the presentation order; - electrode coordinates or anatomical labels. BIDS iEEG requires an electrodes file, so each subject’s electrodes.tsv

lists the released channel names with x, y, z and size set to n/a, and coordsystem.json says Other with no unit. No coordinates were estimated.

Files

  • sub-<id>/ieeg/sub-<id>_task-wordreading_run-<n>_ieeg.vhdr/.vmrk/.eeg: one file per source run (word-recitation/run<n>). BrainVision, IEEE float32, 1000 Hz, unit V, resolution 1. Each value is the authors’ stored float64 value rounded to float32 (maximum absolute rounding error 1.5e-08 V). The epochs are stored back to back: the sidecar says RecordingType: epoched with EpochLength: 3, and the marker file has a New Segment at each epoch start. The time axis of the file is not the time axis of the original recording. Do not filter across epoch boundaries. To get epochs, split every 3000 samples. mne_bids.read_raw_bids refuses epoched recordings, so read the file with mne.io.read_raw_brainvision and reshape it.

  • ..._events.tsv: one row per epoch. value holds the word exactly as stored (UTF-8 Chinese) and word_en holds the authors’ translation. source_epoch_index is the epoch’s position in the source pickle.

  • ..._channels.tsv: all channels are SEEG with the authors’ filter settings. group is the electrode shaft (the name prefix). All channels are good, since the release flags none.

  • sub-<id>_scans.tsv: the source file behind each recording.

  • sourcedata/huggingface-liulab-repository-Du-IN/: the original data and info files of every run, byte-identical (sha256 in sourcedata/sourcedata_provenance.json). They are Python pickles. Never call pickle.load on files you do not trust. code/laneJ_pkl.py is a restricted unpickler that admits only the globals these files use. Each data file is a list of records {name: <word>, data_s: float64 array (n_channels, 3000)}. Each info file holds {ch_names: [...]}. Model checkpoints (pretrains/) from the same release are not copied, because they are not recordings.

  • code/: the download, conversion and round-trip scripts used to build this package.

Recordings

| Subject | Runs | Channels | Epochs | Minutes of epochs | NaN values |
|---|---|---|---|---|---|
| sub-001 | 1, 2, 3, 4, 5, 6 | 104 | 3065 | 153.2 | 0 |
| sub-002 | 1, 2, 3, 4, 5, 6 | 115 | 3625 | 181.2 | 0 |
| sub-003 | 1, 2, 3, 4, 5, 6 | 67 | 3442 | 172.1 | 0 |
| sub-004 | 1, 3, 4, 5, 6 | 100 | 3043 | 152.2 | 0 |
| sub-005 | 1, 2, 3, 4, 5 | 123 | 3048 | 152.4 | 0 |
| sub-006 | 1, 2, 3, 4, 5, 6 | 101 | 3365 | 168.2 | 0 |
| sub-007 | 1, 2, 3, 4, 5, 6 | 94 | 3546 | 177.3 | 0 |
| sub-008 | 1, 2, 3, 4, 5, 6 | 98 | 3628 | 181.4 | 0 |
| sub-009 | 1, 2, 3, 4, 5, 6 | 147 | 3473 | 173.7 | 0 |
| sub-010 | 1, 2, 3, 4, 5, 6 | 88 | 3501 | 175.1 | 0 |
| sub-011 | 1, 2, 3, 4, 5, 6 | 90 | 3161 | 158.1 | 0 |
| sub-012 | 1, 2, 3 | 62 | 1830 | 91.5 | 0 |

Total: 67 runs, 38727 epochs (32.27 h).

Subject 004 has no run 2, subject 005 has no run 6, and subject 012 has only runs 1-3. This matches the release. A full run would hold 610 epochs (5 blocks x 122 trials), but many runs hold fewer. The authors’ loader explains why: “The labels may not be the same, due to bad trials” (utils/data/seeg/he2023xuanwu/word_recitation.py). The removed trials and the rejection criterion are not documented, and sub-009 run 5 has 60 of the 61 words. A few runs contain large transients: the maximum |value| is 0.48 V in sub-003 run 3 and 0.24 V in sub-008 run 4, while median channel SDs are about 2e-5 to 6e-5 V. These are kept as released, and no channel is marked bad.

Conversion and checks

Packaging only: no filtering, resampling, re-referencing, channel dropping or epoch dropping. A round trip read every file back with MNE and compared it with the source pickles. All samples were equal to float32(source), and channel order, epoch count, words, onsets and durations matched in all 67 runs. The BIDS validator reported no errors. Packaged for NEMAR by Bruno Aristimunha (2026-10-06).

How to load

import mne, numpy as np, pandas as pd
vhdr = "sub-001/ieeg/sub-001_task-wordreading_run-1_ieeg.vhdr"
raw = mne.io.read_raw_brainvision(vhdr, preload=True)          # 1000 Hz, SEEG, volts
data = raw.get_data()                                           # (n_channels, n_epochs * 3000)
epochs = data.reshape(data.shape[0], -1, 3000).transpose(1, 0, 2)  # (n_epochs, n_channels, 3000)
ev = pd.read_csv(vhdr.replace("_ieeg.vhdr", "_events.tsv"), sep="\t")
labels = ev["value"].to_numpy()                                 # one Chinese word per epoch (word_en: English)

task-wordreading_events.json lists the 61 words with the authors’ English translations (value Levels).

Funding and acknowledgements

National Science and Technology Innovation 2030 Major Program (2022ZD0205500), National Natural Science Foundation of China (32271093), Beijing Natural Science Foundation (Z230010, L222033), and the Fundamental Research Funds for the Central Universities (paper, Acknowledgements; also in dataset_description.json).

Licence and citation

Licence: CC-BY-4.0 (Hugging Face dataset card of liulab-repository/Du-IN). The Du-IN code on GitHub is MIT licensed.

If you use these data, cite:

Zheng H, Wang H-T, Jiang W-B, Chen Z-T, He L, Lin P-Y, Wei P-H, Zhao G-G, Liu Y-Z. Du-IN: Discrete units-guided mask modeling for decoding speech from Intracranial Neural signals. NeurIPS 2024. doi:10.52202/079017-2542, arXiv:2405.11459.

NEMAR Metadata#

[![DOI](https://img.shields.io/badge/DOI-10.82901%2Fnemar.nm000365-blue)](https://doi.org/10.82901/nemar.nm000365) # Du-IN Chinese word-reading sEEG (Zheng et al., NeurIPS 2024): the authors’ preprocessed epochs These are not raw recordings. This dataset is a BIDS packaging of the word-reading sEEG data released with Du-IN: Discrete units-guided mask modeling for decoding speech from Intracranial Neural signals (Zheng, Wang, Jiang, Chen, He, Lin, Wei, Zhao, Liu; NeurIPS 2024; doi:10.52202/079017-2542; arXiv:2405.11459). The authors published only a preprocessed version of the 61-word reading data on Hugging Face (liulab-repository/Du-IN, revision 2c85b670230dced33f2f1cbd0903648680ecc7d0, last modified 2025-04-30, licence CC-BY-4.0). Their GitHub README says: “we can only share the preprocessed version of the 61-word reading sEEG dataset (~3 hours)”. The ~12 h of non-task recordings used for pre-training are not shared. dataset_description.json therefore declares DatasetType: derivative. ## Participants and task (from the paper, Appendix A) - 12 patients with pharmacologically intractable epilepsy (9 male, 3 female; aged 15-53, mean 27.8, SD 10.4). The release

does not link age or sex to individual subjects, so participants.tsv holds only the subject IDs. 7 to 13 sEEG electrodes per subject, placed for clinical reasons only. Recorded at 2000 Hz.

  • Task: reading Chinese words aloud. The 61-word set and the authors’ English translations are in the paper (Table 4) and in the word_en column of the events files. Each trial starts with the word in white text (0.5 s). The text then turns green, which is the go cue, and stays on screen for 2 s while the subject reads the word aloud. A fixation cross follows for 0.5 s. Each ~10-minute block presents every word twice in random order (122 trials). The paper reports 25 blocks in total.

  • Ethics, as stated by the authors: “Experiments that contribute to this work were approved by IRB. All subjects consent to participate.” The consent form includes a privacy statement and data retention “after deleting all personal identification information (PII)”. The release contains no names, dates or imaging.

## Additional study details (from the paper, arXiv:2405.11459v3: section 4.1, Appendix A, Ethics Statement, Appendix O) - Implant: 7 to 13 sEEG depth electrodes per subject; the paper gives 72 to 158 channels per subject at acquisition

(the released bipolar channel counts are 62 to 147, see the table below). Electrodes carry 8-16 channels each (section 4.3). “All electrode locations are exclusively dictated by clinical considerations.” Most subjects had electrodes on one hemisphere only; the paper shows per-subject electrode locations as figures (Appendix O, Figures 9-11) but no coordinates were released.

  • Recordings: about 15 h of 2000 Hz sEEG per subject, of which about 3 h are task recordings (the part released here) and about 12 h are non-task recordings during wakefulness (not released). The subject’s voice was recorded simultaneously during the task; the audio is not part of the release.

  • Word set: 61 words chosen for (1) versatility in generating sentences, (2) expressing basic caregiving needs and (3) covering as many Chinese pronunciation combinations as possible. The design follows Moses et al. (2021).

  • Recording site: not stated explicitly by the paper or the release. Two authors (P.-H. Wei, G.-G. Zhao) are affiliated with Capital Medical University, Xuanwu Hospital, Beijing, and the release directory is named seeg.he2023xuanwu. No InstitutionName is set in the sidecars for this reason.

  • Amplifier/recording system, electrode manufacturer and acquisition reference are not stated in the paper or the release (left out of the sidecars).

  • Consent (Ethics Statement): adults with full civil capacity signed written informed consent; for minors or participants without full civil capacity, the legal guardian signed. The consent form covers research purpose, risks, data use, a privacy statement (PII not disclosed), data retention after deleting all PII, voluntariness and the right to withdraw at any time.

## Preprocessing done by the authors (before release) Paper section 4.2: band-pass 0.5-200 Hz, 50 Hz notch, resampling to 1000 Hz, bipolar re-reference. The signals were then cut into 3-s samples (3000 points), each labelled with its word. The released variant is named dataset.bipolar.default.unaligned. The paper also lists per-channel z-scoring, but the stored values are not z-scored: per-channel standard deviations are of order 1e-5. That step is presumably applied at training time. The release does not state: - the unit (the magnitudes are only plausible as volts, so V is used and the numbers are left unchanged); - the contact pairing of each bipolar channel (channel names are as released); - where each 3-s epoch sits relative to word onset, go cue or speech onset; - whether the stored epoch order is the presentation order; - electrode coordinates or anatomical labels. BIDS iEEG requires an electrodes file, so each subject’s electrodes.tsv

lists the released channel names with x, y, z and size set to n/a, and coordsystem.json says Other with no unit. No coordinates were estimated.

## Files - sub-<id>/ieeg/sub-<id>_task-wordreading_run-<n>_ieeg.vhdr/.vmrk/.eeg: one file per source run

(word-recitation/run<n>). BrainVision, IEEE float32, 1000 Hz, unit V, resolution 1. Each value is the authors’ stored float64 value rounded to float32 (maximum absolute rounding error 1.5e-08 V). The epochs are stored back to back: the sidecar says RecordingType: epoched with EpochLength: 3, and the marker file has a New Segment at each epoch start. The time axis of the file is not the time axis of the original recording. Do not filter across epoch boundaries. To get epochs, split every 3000 samples. mne_bids.read_raw_bids refuses epoched recordings, so read the file with mne.io.read_raw_brainvision and reshape it.

  • …_events.tsv: one row per epoch. value holds the word exactly as stored (UTF-8 Chinese) and word_en holds the authors’ translation. source_epoch_index is the epoch’s position in the source pickle.

  • …_channels.tsv: all channels are SEEG with the authors’ filter settings. group is the electrode shaft (the name prefix). All channels are good, since the release flags none.

  • sub-<id>_scans.tsv: the source file behind each recording.

  • sourcedata/huggingface-liulab-repository-Du-IN/: the original data and info files of every run, byte-identical (sha256 in sourcedata/sourcedata_provenance.json). They are Python pickles. Never call pickle.load on files you do not trust. code/laneJ_pkl.py is a restricted unpickler that admits only the globals these files use. Each data file is a list of records {name: <word>, data_s: float64 array (n_channels, 3000)}. Each info file holds {ch_names: […]}. Model checkpoints (pretrains/) from the same release are not copied, because they are not recordings.

  • code/: the download, conversion and round-trip scripts used to build this package.

## Recordings | Subject | Runs | Channels | Epochs | Minutes of epochs | NaN values | |---|—|---|—|---|—| | sub-001 | 1, 2, 3, 4, 5, 6 | 104 | 3065 | 153.2 | 0 | | sub-002 | 1, 2, 3, 4, 5, 6 | 115 | 3625 | 181.2 | 0 | | sub-003 | 1, 2, 3, 4, 5, 6 | 67 | 3442 | 172.1 | 0 | | sub-004 | 1, 3, 4, 5, 6 | 100 | 3043 | 152.2 | 0 | | sub-005 | 1, 2, 3, 4, 5 | 123 | 3048 | 152.4 | 0 | | sub-006 | 1, 2, 3, 4, 5, 6 | 101 | 3365 | 168.2 | 0 | | sub-007 | 1, 2, 3, 4, 5, 6 | 94 | 3546 | 177.3 | 0 | | sub-008 | 1, 2, 3, 4, 5, 6 | 98 | 3628 | 181.4 | 0 | | sub-009 | 1, 2, 3, 4, 5, 6 | 147 | 3473 | 173.7 | 0 | | sub-010 | 1, 2, 3, 4, 5, 6 | 88 | 3501 | 175.1 | 0 | | sub-011 | 1, 2, 3, 4, 5, 6 | 90 | 3161 | 158.1 | 0 | | sub-012 | 1, 2, 3 | 62 | 1830 | 91.5 | 0 | Total: 67 runs, 38727 epochs (32.27 h). Subject 004 has no run 2, subject 005 has no run 6, and subject 012 has only runs 1-3. This matches the release. A full run would hold 610 epochs (5 blocks x 122 trials), but many runs hold fewer. The authors’ loader explains why: “The labels may not be the same, due to bad trials” (utils/data/seeg/he2023xuanwu/word_recitation.py). The removed trials and the rejection criterion are not documented, and sub-009 run 5 has 60 of the 61 words. A few runs contain large transients: the maximum |value| is 0.48 V in sub-003 run 3 and 0.24 V in sub-008 run 4, while median channel SDs are about 2e-5 to 6e-5 V. These are kept as released, and no channel is marked bad. ## Conversion and checks Packaging only: no filtering, resampling, re-referencing, channel dropping or epoch dropping. A round trip read every file back with MNE and compared it with the source pickles. All samples were equal to float32(source), and channel order, epoch count, words, onsets and durations matched in all 67 runs. The BIDS validator reported no errors. Packaged for NEMAR by Bruno Aristimunha (2026-10-06). ## How to load `python import mne, numpy as np, pandas as pd vhdr = "sub-001/ieeg/sub-001_task-wordreading_run-1_ieeg.vhdr" raw = mne.io.read_raw_brainvision(vhdr, preload=True)          # 1000 Hz, SEEG, volts data = raw.get_data()                                           # (n_channels, n_epochs * 3000) epochs = data.reshape(data.shape[0], -1, 3000).transpose(1, 0, 2)  # (n_epochs, n_channels, 3000) ev = pd.read_csv(vhdr.replace("_ieeg.vhdr", "_events.tsv"), sep="\t") labels = ev["value"].to_numpy()                                 # one Chinese word per epoch (word_en: English) ` task-wordreading_events.json lists the 61 words with the authors’ English translations (value Levels). ## Funding and acknowledgements National Science and Technology Innovation 2030 Major Program (2022ZD0205500), National Natural Science Foundation of China (32271093), Beijing Natural Science Foundation (Z230010, L222033), and the Fundamental Research Funds for the Central Universities (paper, Acknowledgements; also in dataset_description.json). ## Licence and citation Licence: CC-BY-4.0 (Hugging Face dataset card of liulab-repository/Du-IN). The Du-IN code on GitHub is MIT licensed. If you use these data, cite: Zheng H, Wang H-T, Jiang W-B, Chen Z-T, He L, Lin P-Y, Wei P-H, Zhao G-G, Liu Y-Z. Du-IN: Discrete units-guided mask modeling for decoding speech from Intracranial Neural signals. NeurIPS 2024. doi:10.52202/079017-2542, arXiv:2405.11459.

License: CC-BY-4.0

Authors:

  • Hui Zheng

  • Hai-Teng Wang

  • Wei-Bang Jiang

  • Zhong-Tao Chen

  • Li He

  • … and 4 more

Versions:

Version

DOI

Released

current

10.82901/nemar.nm000365

§ 03Cohort · Participants

Cohort#

Dataset Statistics#

Channel counts (ch)

626788909498100101104115123147

Sampling frequencies: 1000.0 Hz (n=67 recordings)

Total recording duration: 32 h

§ 04Signal · Electrodes & trace

Signal · Electrodes & live trace#

Fig. 01 Signal & montage 90 (6), 94 (6), 115 (6), 147 (6), 67 (6), 104 (6), 88 (6), 101 (6), 98 (6), 123 (5), 100 (5), 62 (3) ch · iEEG · 1000 Hz · 12 subjects, 67 recordings
Live trace viewer — sub-012 · task-wordreading · run-3

Showing one representative recording out of 12 subjects and 67 recordings in this dataset. Browse the full set on OpenNeuro; drop any other _ieeg.{set,edf,bdf,vhdr} file onto the viewer (or pass ?ieeg=<url>) to inspect it.

No scalp electrode layout is currently indexed for this dataset. Once the eegdash montage registry ingests it, the interactive viewer will appear here automatically.

NEMAR Processing Statistics#

The plots below are generated by NEMAR’s automated EEG pipeline. The histogram shows pipeline success for data cleaning and ICA decomposition, the percentage of data frames and EEG channels retained after artefact removal, line noise per channel (RMS, dB), and the age/gender distribution of participants.

HED event descriptors word cloud HED event descriptors word cloud — NM000365
§ 05Manifest · BIDS tree

Manifest#

File Explorer#

Browse the BIDS file structure of this dataset. Records are fetched on demand from the EEGDash catalog the first time you open the explorer.

Recordings—
Files—
Subjects—
Modalities—
Click to load file structure…
Full dataset metadata table

Dataset ID

NM000365

Title

Du-IN Chinese word-reading sEEG (Zheng et al., NeurIPS 2024): authors’ preprocessed 3-s epochs, 12 subjects, 61 words

Author (year)

—

Canonical

—

Importable as

NM000365

Year

2021

Authors

Hui Zheng, Hai-Teng Wang, Wei-Bang Jiang, Zhong-Tao Chen, Li He, Pei-Yang Lin, Peng-Hu Wei, Guo-Guang Zhao, Yun-Zhe Liu

License

CC-BY-4.0

Citation / DOI

10.82901/nemar.nm000365

Source links

OpenNeuro | NeMAR | Source URL

Copy-paste BibTeX
@dataset{nm000365,
  title = {Du-IN Chinese word-reading sEEG (Zheng et al., NeurIPS 2024): authors' preprocessed 3-s epochs, 12 subjects, 61 words},
  author = {Hui Zheng and Hai-Teng Wang and Wei-Bang Jiang and Zhong-Tao Chen and Li He and Pei-Yang Lin and Peng-Hu Wei and Guo-Guang Zhao and Yun-Zhe Liu},
  doi = {10.82901/nemar.nm000365},
  url = {https://doi.org/10.82901/nemar.nm000365},
}
§ 06API · Programmatic access

API Reference#

Signature
eegdash.dataset
class
eegdash.dataset.NM000365(cache_dir, query=None, s3_bucket=None, **kwargs)
Bases: EEGDashDataset
Author (year)—
Canonical—
Importable asNM000365
Sourceeegdash/dataset/registry.py · [source ↗]
class eegdash.dataset.NM000365(cache_dir: str, query: dict | None = None, s3_bucket: str | None = None, **kwargs)[source]#

Du-IN Chinese word-reading sEEG (Zheng et al., NeurIPS 2024): authors’ preprocessed 3-s epochs, 12 subjects, 61 words

Study:

nm000365 (NeMAR)

Author (year):

—

Canonical:

—

Also importable as: NM000365.

Modality: ieeg; Subject type: Unknown. Subjects: 12; recordings: 67; tasks: 1.

Parameters:
  • cache_dir (str | Path) – Directory where data are cached locally.

  • query (dict | None) – Additional MongoDB-style filters to AND with the dataset selection. Must not contain the key dataset.

  • s3_bucket (str | None) – Base S3 bucket used to locate the data.

  • **kwargs (dict) – Additional keyword arguments forwarded to EEGDashDataset.

data_dir#

Local dataset cache directory (cache_dir / dataset_id).

Type:

Path

query#

Merged query with the dataset filter applied.

Type:

dict

records#

Metadata records used to build the dataset, if pre-fetched.

Type:

list[dict] | None

Notes

Each item is a recording; recording-level metadata are available via dataset.description. query supports MongoDB-style filters on fields in ALLOWED_QUERY_FIELDS and is combined with the dataset filter. Dataset-specific caveats are not provided in the summary metadata.

References

OpenNeuro dataset: https://openneuro.org/datasets/nm000365 NeMAR dataset: https://nemar.org/dataexplorer/detail?dataset_id=nm000365 DOI: https://doi.org/10.82901/nemar.nm000365

Examples

>>> from eegdash.dataset import NM000365
>>> dataset = NM000365(cache_dir="./data")
>>> recording = dataset[0]
>>> raw = recording.load()
__init__(cache_dir: str, query: dict | None = None, s3_bucket: str | None = None, **kwargs)[source]#
save(path: str, overwrite: bool = False, offset: int = 0)[source]#

Save datasets to files by creating one subdirectory for each dataset:

path/
    0/
        0-raw.fif | 0-epo.fif
        description.json
        raw_preproc_kwargs.json (if raws were preprocessed)
        window_kwargs.json (if this is a windowed dataset)
        window_preproc_kwargs.json  (if windows were preprocessed)
        target_name.json (if target_name is not None and dataset is raw)
    1/
        1-raw.fif | 1-epo.fif
        description.json
        raw_preproc_kwargs.json (if raws were preprocessed)
        window_kwargs.json (if this is a windowed dataset)
        window_preproc_kwargs.json  (if windows were preprocessed)
        target_name.json (if target_name is not None and dataset is raw)
Parameters:
  • path (str) –

    Directory in which subdirectories are created to store

    -raw.fif | -epo.fif and .json files to.

  • overwrite (bool) – Whether to delete old subdirectories that will be saved to in this call.

  • offset (int) – If provided, the integer is added to the id of the dataset in the concat. This is useful in the setting of very large datasets, where one dataset has to be processed and saved at a time to account for its original position.

Access modesMNE → braindecode → PyTorch → ML
.rawMNE Raw object — standard tools (filter, epoch, ICA, plot_psd).mne
DataLoaderWraps the windowed dataset into a PyTorch DataLoader; supports parallel workers and on-the-fly augmentations.pytorch
Zarr cacheOptional braindecode Zarr mirror for fast resume; persisted to cache_dir.zarr
Hugging FaceNo per-dataset mirror published yet — browse the EEGDash org listing for sibling datasets. See the datasets loader API.huggingface
Croissant 1.0Machine-readable JSON-LD descriptor — NM000365.croissant.json (MLCommons schema, ingestible by PyTorch / TensorFlow / JAX).mlcommons
Examples using EEGDashcurated · start here

Swap any load_dataset(...) call for nm000365 to reproduce the tutorial on this dataset.

Citation

Hui Zheng, Hai-Teng Wang, Wei-Bang Jiang, Zhong-Tao Chen, Li He, … (2021). Du-IN Chinese word-reading sEEG (Zheng et al., NeurIPS 2024): authors' preprocessed 3-s epochs, 12 subjects, 61 words. 10.82901/nemar.nm000365

Provenance

¹Contributed to nemar in BIDS format.

²Curated & ingested by the EEGDash catalog; see CITATION.cff for canonical reference.

³Persistent identifier: 10.82901/nemar.nm000365.

BIDS
BIDS 1.10.0
Sidecars
events · channels · electrodes · coordsystem · eeg.json
Provenance
Machine-readable
Mirrors

See Also#