NM000111: eeg dataset, 118 subjects#
ISRUC-Sleep: A comprehensive public dataset for sleep researchers.
Access recordings and metadata through EEGDash.
Citation: Sirvan Khalighi, Teresa Sousa, Jose Moutinho Santos, Urbano Nunes (2016). ISRUC-Sleep: A comprehensive public dataset for sleep researchers.. 10.82901/nemar.nm000111
Modality: eeg Subjects: 118 Recordings: 126 License: n/a Source: nemar
Metadata: Complete (100%)
118-participant EEG dataset — ISRUC-Sleep: A comprehensive public dataset for sleep researchers..
Quickstart#
Install
pip install eegdash
Access the data
from eegdash.dataset import NM000111
dataset = NM000111(cache_dir="./data")
# Get the raw object of the first recording
raw = dataset.datasets[0].raw
print(raw.info)
Filter by subject
dataset = NM000111(cache_dir="./data", subject="01")
Advanced query
dataset = NM000111(
cache_dir="./data",
query={"subject": {"$in": ["01", "02"]}},
)
Iterate recordings
for rec in dataset:
print(rec.subject, rec.raw.info['sfreq'])
If you use this dataset in your research, please cite the original authors.
BibTeX
@dataset{nm000111,
title = {ISRUC-Sleep: A comprehensive public dataset for sleep researchers.},
author = {Sirvan Khalighi and Teresa Sousa and Jose Moutinho Santos and Urbano Nunes},
doi = {10.82901/nemar.nm000111},
url = {https://doi.org/10.82901/nemar.nm000111},
}
About This Dataset#
The ISRUC-Sleep dataset comprises overnight polysomnographic (PSG) recordings and manual sleep stage annotations across three subgroups. The data support research in automatic sleep staging and sleep-disordered breathing. Signals include EEG, EOG, EMG, respiratory channels and others, provided as EDF-compatible .rec files. For each recording, sleep was scored by two expert scorers in 30-second epochs.
Participants slept overnight in a clinical environment with standard PSG montage. Two independent human scorers labeled each 30-second epoch into sleep stages following AASM/R&K guidelines used by the dataset (W, N1, N2, N3, and REM). The dataset is divided into subgroups with different focuses (e.g., subjects with sleep disorders, multiple nights). Please refer to the publication and the Details spreadsheets for demographic and clinical descriptors.
ISRUC-Sleep
Introduction
Description of the preprocessing if any
Original .rec files are symlinked (or copied if needed) to .edf without modification. Sleep stages come from the scorer-1 Excel files (col0 epoch, col1 label; headers auto-skipped; NaN epochs filled sequentially; unknown labels -> U). Labels from the second scorer, when present, are stored in annotation extras. Measurement dates use Date of recording from Details (UTC) when available, otherwise 2020-01-01. Participant demographics (Sex, Age) are pulled directly from the Details spreadsheets.
View full README
ISRUC-Sleep
Introduction
Description of the preprocessing if any
Original .rec files are symlinked (or copied if needed) to .edf without modification. Sleep stages come from the scorer-1 Excel files (col0 epoch, col1 label; headers auto-skipped; NaN epochs filled sequentially; unknown labels -> U). Labels from the second scorer, when present, are stored in annotation extras. Measurement dates use Date of recording from Details (UTC) when available, otherwise 2020-01-01. Participant demographics (Sex, Age) are pulled directly from the Details spreadsheets.
Description of the event values
Sleep stages are encoded per 30 s epoch. The following mapping is used: - 0: Sleep stage W (Wake) - 1: Sleep stage N1 - 2: Sleep stage N2 - 3: Sleep stage N3 - 5: Sleep stage R (REM) - 6: Sleep stage U (Unknown)
The annotations are added as events with onset at the epoch start, duration 30 seconds, and description matching the above labels.
Citation
When using this dataset, please cite: 1. Khalighi S., Sousa T., Santos J.M., Nunes U. ISRUC-Sleep: A comprehensive public dataset for sleep researchers.
Computer methods and programs in biomedicine 124 (2016): 180-192. DOI: 10.1016/j.cmpb.2015.10.013
Project site: https://sleeptight.isr.uc.pt
Data curators (BIDS conversion): Pierre Guetschel Data collectors (original dataset):
Sirvan Khalighi; Teresa Sousa; Jose Moutinho Santos; Urbano Nunes
Automatic report
Report automatically generated by ``mne_bids.make_report()``.
The ISRUC-Sleep dataset was created by Sirvan Khalighi, Teresa Sousa, Jose
Moutinho Santos, and Urbano Nunes and conforms to BIDS version 1.7.0. This report was generated with MNE-BIDS (https://doi.org/10.21105/joss.01896). The dataset consists of 118 participants (comprised of 71 male and 47 female participants; handedness were all unknown; ages all unknown) and 2 recording sessions: 1, and 2. Data was recorded using an EEG system sampled at 200.0 Hz with line noise at n/a Hz. There were 126 scans in total. Recording durations ranged from 10724.0 to 31860.0 seconds (mean = 26751.84, std = 2534.94), for a total of 3370731.37 seconds of data recorded over all scans. For each dataset, there were on average 19.63 (std = 0.65) recording channels per scan, out of which 19.63 (std = 0.65) were used in analysis (0.0 +/- 0.0 were removed from analysis).
NEMAR Metadata#
[](https://doi.org/10.82901/nemar.nm000111) # ISRUC-Sleep ## Introduction The ISRUC-Sleep dataset comprises overnight polysomnographic (PSG) recordings and manual sleep stage annotations across three subgroups. The data support research in automatic sleep staging and sleep-disordered breathing. Signals include EEG, EOG, EMG, respiratory channels and others, provided as EDF-compatible .rec files. For each recording, sleep was scored by two expert scorers in 30-second epochs. ## Overview of the experiment Participants slept overnight in a clinical environment with standard PSG montage. Two independent human scorers labeled each 30-second epoch into sleep stages following AASM/R&K guidelines used by the dataset (W, N1, N2, N3, and REM). The dataset is divided into subgroups with different focuses (e.g., subjects with sleep disorders, multiple nights). Please refer to the publication and the Details spreadsheets for demographic and clinical descriptors. ## Description of the preprocessing if any Original .rec files are symlinked (or copied if needed) to .edf without modification. Sleep stages come from the scorer-1 Excel files (col0 epoch, col1 label; headers auto-skipped; NaN epochs filled sequentially; unknown labels -> U). Labels from the second scorer, when present, are stored in annotation extras. Measurement dates use Date of recording from Details (UTC) when available, otherwise 2020-01-01. Participant demographics (Sex, Age) are pulled directly from the Details spreadsheets. ## Description of the event values Sleep stages are encoded per 30 s epoch. The following mapping is used: - 0: Sleep stage W (Wake) - 1: Sleep stage N1 - 2: Sleep stage N2 - 3: Sleep stage N3 - 5: Sleep stage R (REM) - 6: Sleep stage U (Unknown) The annotations are added as events with onset at the epoch start, duration 30 seconds, and description matching the above labels. ## Citation When using this dataset, please cite: 1. Khalighi S., Sousa T., Santos J.M., Nunes U. ISRUC-Sleep: A comprehensive public dataset for sleep researchers.
Computer methods and programs in biomedicine 124 (2016): 180-192. DOI: 10.1016/j.cmpb.2015.10.013
2. Project site: https://sleeptight.isr.uc.pt Data curators (BIDS conversion): Pierre Guetschel Data collectors (original dataset): Sirvan Khalighi; Teresa Sousa; Jose Moutinho Santos; Urbano Nunes — ## Automatic report Report automatically generated by `mne_bids.make_report()`. > The ISRUC-Sleep dataset was created by Sirvan Khalighi, Teresa Sousa, Jose Moutinho Santos, and Urbano Nunes and conforms to BIDS version 1.7.0. This report was generated with MNE-BIDS (https://doi.org/10.21105/joss.01896). The dataset consists of 118 participants (comprised of 71 male and 47 female participants; handedness were all unknown; ages all unknown) and 2 recording sessions: 1, and 2. Data was recorded using an EEG system sampled at 200.0 Hz with line noise at n/a Hz. There were 126 scans in total. Recording durations ranged from 10724.0 to 31860.0 seconds (mean = 26751.84, std = 2534.94), for a total of 3370731.37 seconds of data recorded over all scans. For each dataset, there were on average 19.63 (std = 0.65) recording channels per scan, out of which 19.63 (std = 0.65) were used in analysis (0.0 +/- 0.0 were removed from analysis).
License: n/a
Authors:
Sirvan Khalighi
Teresa Sousa
Jose Moutinho Santos
Urbano Nunes
Versions:
Version |
DOI |
Released |
|---|---|---|
|
Cohort#
Dataset Statistics#
Age distribution by gender (n=116, range 20–85 yr, mean 49.5 yr)
Sex composition
Channel counts (ch)
Sampling frequencies: 200.0 Hz (n=126 recordings)
Total recording duration: 936 h
Signal · Electrodes & live trace#
Live trace viewer — sub-I090 · task-sleep
Showing one representative recording out of
118 subjects and 126 recordings in this dataset.
Browse the full set on OpenNeuro;
drop any other _eeg.{set,edf,bdf,vhdr} file onto the
viewer (or pass ?eeg=<url>) to inspect it.
No scalp electrode layout is currently indexed for this dataset. Once the eegdash montage registry ingests it, the interactive viewer will appear here automatically.
NEMAR Processing Statistics#
The plots below are generated by NEMAR’s automated EEG pipeline. The histogram shows pipeline success for data cleaning and ICA decomposition, the percentage of data frames and EEG channels retained after artefact removal, line noise per channel (RMS, dB), and the age/gender distribution of participants.
HED event descriptors word cloud
Manifest#
File Explorer#
Browse the BIDS file structure of this dataset. Records are fetched on demand from the EEGDash catalog the first time you open the explorer.
Full dataset metadata table
Dataset ID |
|
Title |
ISRUC-Sleep: A comprehensive public dataset for sleep researchers. |
Author (year) |
— |
Canonical |
— |
Importable as |
|
Year |
2016 |
Authors |
Sirvan Khalighi, Teresa Sousa, Jose Moutinho Santos, Urbano Nunes |
License |
n/a |
Citation / DOI |
|
Source links |
OpenNeuro | NeMAR | Source URL |
Copy-paste BibTeX
@dataset{nm000111,
title = {ISRUC-Sleep: A comprehensive public dataset for sleep researchers.},
author = {Sirvan Khalighi and Teresa Sousa and Jose Moutinho Santos and Urbano Nunes},
doi = {10.82901/nemar.nm000111},
url = {https://doi.org/10.82901/nemar.nm000111},
}
API Reference#
eegdash.datasetEEGDashDataset- class eegdash.dataset.NM000111(cache_dir: str, query: dict | None = None, s3_bucket: str | None = None, **kwargs)[source]#
ISRUC-Sleep: A comprehensive public dataset for sleep researchers.
- Study:
nm000111(NeMAR)- Author (year):
—
- Canonical:
—
Also importable as:
NM000111.Modality:
eeg; Subject type:Unknown. Subjects: 118; recordings: 126; tasks: 1.- Parameters:
cache_dir (str | Path) – Directory where data are cached locally.
query (dict | None) – Additional MongoDB-style filters to AND with the dataset selection. Must not contain the key
dataset.s3_bucket (str | None) – Base S3 bucket used to locate the data.
**kwargs (dict) – Additional keyword arguments forwarded to
EEGDashDataset.
- data_dir#
Local dataset cache directory (
cache_dir / dataset_id).- Type:
Path
Notes
Each item is a recording; recording-level metadata are available via
dataset.description.querysupports MongoDB-style filters on fields inALLOWED_QUERY_FIELDSand is combined with the dataset filter. Dataset-specific caveats are not provided in the summary metadata.References
OpenNeuro dataset: https://openneuro.org/datasets/nm000111 NeMAR dataset: https://nemar.org/dataexplorer/detail?dataset_id=nm000111 DOI: https://doi.org/10.82901/nemar.nm000111
Examples
>>> from eegdash.dataset import NM000111 >>> dataset = NM000111(cache_dir="./data") >>> recording = dataset[0] >>> raw = recording.load()
- __init__(cache_dir: str, query: dict | None = None, s3_bucket: str | None = None, **kwargs)[source]#
- save(path: str, overwrite: bool = False, offset: int = 0)[source]#
Save datasets to files by creating one subdirectory for each dataset:
path/ 0/ 0-raw.fif | 0-epo.fif description.json raw_preproc_kwargs.json (if raws were preprocessed) window_kwargs.json (if this is a windowed dataset) window_preproc_kwargs.json (if windows were preprocessed) target_name.json (if target_name is not None and dataset is raw) 1/ 1-raw.fif | 1-epo.fif description.json raw_preproc_kwargs.json (if raws were preprocessed) window_kwargs.json (if this is a windowed dataset) window_preproc_kwargs.json (if windows were preprocessed) target_name.json (if target_name is not None and dataset is raw)
- Parameters:
path (str) –
- Directory in which subdirectories are created to store
-raw.fif | -epo.fif and .json files to.
overwrite (bool) – Whether to delete old subdirectories that will be saved to in this call.
offset (int) – If provided, the integer is added to the id of the dataset in the concat. This is useful in the setting of very large datasets, where one dataset has to be processed and saved at a time to account for its original position.
BaseDataset from braindecode — windowed via create_windows_from_events.braindecodeDataLoader; supports parallel workers and on-the-fly augmentations.pytorchSwap any load_dataset(...) call for nm000111 to reproduce the tutorial on this dataset.
Citation
Sirvan Khalighi, Teresa Sousa, Jose Moutinho Santos, Urbano Nunes (2016). ISRUC-Sleep: A comprehensive public dataset for sleep researchers.. 10.82901/nemar.nm000111
Provenance
¹Contributed to nemar in BIDS format.
²Curated & ingested by the EEGDash catalog; see CITATION.cff for canonical reference.
³Persistent identifier: 10.82901/nemar.nm000111.
See Also#
eegdash.dataset.EEGDashDataseteegdash.dataset