EEGdashNeMARNM000111
Iss. 111 · 118 subjects · 126 recordings · n/a
Dataset Brief · ISRUC-Sleep

NM000111: eeg dataset, 118 subjects#

ISRUC-Sleep: A comprehensive public dataset for sleep researchers.

Access recordings and metadata through EEGDash.

Citation: Sirvan Khalighi, Teresa Sousa, Jose Moutinho Santos, Urbano Nunes (2016). ISRUC-Sleep: A comprehensive public dataset for sleep researchers.. 10.82901/nemar.nm000111

Modality: eeg Subjects: 118 Recordings: 126 License: n/a Source: nemar

Metadata: Complete (100%)

118-participant EEG dataset — ISRUC-Sleep: A comprehensive public dataset for sleep researchers..

EEG · 20 (70), 19 (53), 21 (2), 24 ch200 HzBIDS 1.7.0Task · sleep2 sessions
Layer 01Study
What was asked
Hypothesis, independent & dependent variables, paradigm, cohort, and the editorial caveats around what the recordings can and cannot answer.
Layer 02Signal · BIDS
What was recorded
Sidecars, channels & electrodes, coordinate system, event semantics, and quality stats from the NEMAR pipeline when available.
Layer 03Training · ML
What you can train on
Recommended access modes — MNE Raw, braindecode windows, PyTorch DataLoader — plus the targets the metadata makes addressable.
§ 01Access · Get started

Quickstart#

Install

pip install eegdash

Access the data

from eegdash.dataset import NM000111

dataset = NM000111(cache_dir="./data")
# Get the raw object of the first recording
raw = dataset.datasets[0].raw
print(raw.info)

Filter by subject

dataset = NM000111(cache_dir="./data", subject="01")

Advanced query

dataset = NM000111(
    cache_dir="./data",
    query={"subject": {"$in": ["01", "02"]}},
)

Iterate recordings

for rec in dataset:
    print(rec.subject, rec.raw.info['sfreq'])

If you use this dataset in your research, please cite the original authors.

BibTeX

@dataset{nm000111,
  title = {ISRUC-Sleep: A comprehensive public dataset for sleep researchers.},
  author = {Sirvan Khalighi and Teresa Sousa and Jose Moutinho Santos and Urbano Nunes},
  doi = {10.82901/nemar.nm000111},
  url = {https://doi.org/10.82901/nemar.nm000111},
}
§ 02Study · The README

About This Dataset#

The ISRUC-Sleep dataset comprises overnight polysomnographic (PSG) recordings and manual sleep stage annotations across three subgroups. The data support research in automatic sleep staging and sleep-disordered breathing. Signals include EEG, EOG, EMG, respiratory channels and others, provided as EDF-compatible .rec files. For each recording, sleep was scored by two expert scorers in 30-second epochs.

Participants slept overnight in a clinical environment with standard PSG montage. Two independent human scorers labeled each 30-second epoch into sleep stages following AASM/R&K guidelines used by the dataset (W, N1, N2, N3, and REM). The dataset is divided into subgroups with different focuses (e.g., subjects with sleep disorders, multiple nights). Please refer to the publication and the Details spreadsheets for demographic and clinical descriptors.

DOI

ISRUC-Sleep

Introduction

Description of the preprocessing if any

Original .rec files are symlinked (or copied if needed) to .edf without modification. Sleep stages come from the scorer-1 Excel files (col0 epoch, col1 label; headers auto-skipped; NaN epochs filled sequentially; unknown labels -> U). Labels from the second scorer, when present, are stored in annotation extras. Measurement dates use Date of recording from Details (UTC) when available, otherwise 2020-01-01. Participant demographics (Sex, Age) are pulled directly from the Details spreadsheets.

View full README

DOI

ISRUC-Sleep

Introduction

Description of the preprocessing if any

Original .rec files are symlinked (or copied if needed) to .edf without modification. Sleep stages come from the scorer-1 Excel files (col0 epoch, col1 label; headers auto-skipped; NaN epochs filled sequentially; unknown labels -> U). Labels from the second scorer, when present, are stored in annotation extras. Measurement dates use Date of recording from Details (UTC) when available, otherwise 2020-01-01. Participant demographics (Sex, Age) are pulled directly from the Details spreadsheets.

Description of the event values

Sleep stages are encoded per 30 s epoch. The following mapping is used: - 0: Sleep stage W (Wake) - 1: Sleep stage N1 - 2: Sleep stage N2 - 3: Sleep stage N3 - 5: Sleep stage R (REM) - 6: Sleep stage U (Unknown)

The annotations are added as events with onset at the epoch start, duration 30 seconds, and description matching the above labels.

Citation

When using this dataset, please cite: 1. Khalighi S., Sousa T., Santos J.M., Nunes U. ISRUC-Sleep: A comprehensive public dataset for sleep researchers.

Computer methods and programs in biomedicine 124 (2016): 180-192. DOI: 10.1016/j.cmpb.2015.10.013

  1. Project site: https://sleeptight.isr.uc.pt

Data curators (BIDS conversion): Pierre Guetschel Data collectors (original dataset):

Sirvan Khalighi; Teresa Sousa; Jose Moutinho Santos; Urbano Nunes

Automatic report

Report automatically generated by ``mne_bids.make_report()``.

The ISRUC-Sleep dataset was created by Sirvan Khalighi, Teresa Sousa, Jose

Moutinho Santos, and Urbano Nunes and conforms to BIDS version 1.7.0. This report was generated with MNE-BIDS (https://doi.org/10.21105/joss.01896). The dataset consists of 118 participants (comprised of 71 male and 47 female participants; handedness were all unknown; ages all unknown) and 2 recording sessions: 1, and 2. Data was recorded using an EEG system sampled at 200.0 Hz with line noise at n/a Hz. There were 126 scans in total. Recording durations ranged from 10724.0 to 31860.0 seconds (mean = 26751.84, std = 2534.94), for a total of 3370731.37 seconds of data recorded over all scans. For each dataset, there were on average 19.63 (std = 0.65) recording channels per scan, out of which 19.63 (std = 0.65) were used in analysis (0.0 +/- 0.0 were removed from analysis).

NEMAR Metadata#

[![DOI](https://img.shields.io/badge/DOI-10.82901%2Fnemar.nm000111-blue)](https://doi.org/10.82901/nemar.nm000111) # ISRUC-Sleep ## Introduction The ISRUC-Sleep dataset comprises overnight polysomnographic (PSG) recordings and manual sleep stage annotations across three subgroups. The data support research in automatic sleep staging and sleep-disordered breathing. Signals include EEG, EOG, EMG, respiratory channels and others, provided as EDF-compatible .rec files. For each recording, sleep was scored by two expert scorers in 30-second epochs. ## Overview of the experiment Participants slept overnight in a clinical environment with standard PSG montage. Two independent human scorers labeled each 30-second epoch into sleep stages following AASM/R&K guidelines used by the dataset (W, N1, N2, N3, and REM). The dataset is divided into subgroups with different focuses (e.g., subjects with sleep disorders, multiple nights). Please refer to the publication and the Details spreadsheets for demographic and clinical descriptors. ## Description of the preprocessing if any Original .rec files are symlinked (or copied if needed) to .edf without modification. Sleep stages come from the scorer-1 Excel files (col0 epoch, col1 label; headers auto-skipped; NaN epochs filled sequentially; unknown labels -> U). Labels from the second scorer, when present, are stored in annotation extras. Measurement dates use Date of recording from Details (UTC) when available, otherwise 2020-01-01. Participant demographics (Sex, Age) are pulled directly from the Details spreadsheets. ## Description of the event values Sleep stages are encoded per 30 s epoch. The following mapping is used: - 0: Sleep stage W (Wake) - 1: Sleep stage N1 - 2: Sleep stage N2 - 3: Sleep stage N3 - 5: Sleep stage R (REM) - 6: Sleep stage U (Unknown) The annotations are added as events with onset at the epoch start, duration 30 seconds, and description matching the above labels. ## Citation When using this dataset, please cite: 1. Khalighi S., Sousa T., Santos J.M., Nunes U. ISRUC-Sleep: A comprehensive public dataset for sleep researchers.

Computer methods and programs in biomedicine 124 (2016): 180-192. DOI: 10.1016/j.cmpb.2015.10.013

2. Project site: https://sleeptight.isr.uc.pt Data curators (BIDS conversion): Pierre Guetschel Data collectors (original dataset): Sirvan Khalighi; Teresa Sousa; Jose Moutinho Santos; Urbano Nunes — ## Automatic report Report automatically generated by `mne_bids.make_report()`. > The ISRUC-Sleep dataset was created by Sirvan Khalighi, Teresa Sousa, Jose Moutinho Santos, and Urbano Nunes and conforms to BIDS version 1.7.0. This report was generated with MNE-BIDS (https://doi.org/10.21105/joss.01896). The dataset consists of 118 participants (comprised of 71 male and 47 female participants; handedness were all unknown; ages all unknown) and 2 recording sessions: 1, and 2. Data was recorded using an EEG system sampled at 200.0 Hz with line noise at n/a Hz. There were 126 scans in total. Recording durations ranged from 10724.0 to 31860.0 seconds (mean = 26751.84, std = 2534.94), for a total of 3370731.37 seconds of data recorded over all scans. For each dataset, there were on average 19.63 (std = 0.65) recording channels per scan, out of which 19.63 (std = 0.65) were used in analysis (0.0 +/- 0.0 were removed from analysis).

License: n/a

Authors:

  • Sirvan Khalighi

  • Teresa Sousa

  • Jose Moutinho Santos

  • Urbano Nunes

Versions:

Version

DOI

Released

current

10.82901/nemar.nm000111

§ 03Cohort · Participants

Cohort#

Dataset Statistics#

Age distribution by gender (n=116, range 20–85 yr, mean 49.5 yr)

2025303540455055606570758085
Female · 47Male · 69

Sex composition

118
subjects
Female
47
Male
71
F : M ratio
0.66 : 1
40% female · n = 118 subjects with reported sex.

Channel counts (ch)

19202124

Sampling frequencies: 200.0 Hz (n=126 recordings)

Total recording duration: 936 h

§ 04Signal · Electrodes & trace

Signal · Electrodes & live trace#

Fig. 01 Signal & montage 20 (70), 19 (53), 21 (2), 24 ch · EEG · 200 Hz · 118 subjects, 126 recordings
Live trace viewer — sub-I090 · task-sleep

Showing one representative recording out of 118 subjects and 126 recordings in this dataset. Browse the full set on OpenNeuro; drop any other _eeg.{set,edf,bdf,vhdr} file onto the viewer (or pass ?eeg=<url>) to inspect it.

No scalp electrode layout is currently indexed for this dataset. Once the eegdash montage registry ingests it, the interactive viewer will appear here automatically.

NEMAR Processing Statistics#

The plots below are generated by NEMAR’s automated EEG pipeline. The histogram shows pipeline success for data cleaning and ICA decomposition, the percentage of data frames and EEG channels retained after artefact removal, line noise per channel (RMS, dB), and the age/gender distribution of participants.

HED event descriptors word cloud HED event descriptors word cloud — NM000111
§ 05Manifest · BIDS tree

Manifest#

File Explorer#

Browse the BIDS file structure of this dataset. Records are fetched on demand from the EEGDash catalog the first time you open the explorer.

Recordings
Files
Subjects
Modalities
Click to load file structure…
Full dataset metadata table

Dataset ID

NM000111

Title

ISRUC-Sleep: A comprehensive public dataset for sleep researchers.

Author (year)

Canonical

Importable as

NM000111

Year

2016

Authors

Sirvan Khalighi, Teresa Sousa, Jose Moutinho Santos, Urbano Nunes

License

n/a

Citation / DOI

10.82901/nemar.nm000111

Source links

OpenNeuro | NeMAR | Source URL

Copy-paste BibTeX
@dataset{nm000111,
  title = {ISRUC-Sleep: A comprehensive public dataset for sleep researchers.},
  author = {Sirvan Khalighi and Teresa Sousa and Jose Moutinho Santos and Urbano Nunes},
  doi = {10.82901/nemar.nm000111},
  url = {https://doi.org/10.82901/nemar.nm000111},
}
§ 06API · Programmatic access

API Reference#

Signature
eegdash.dataset
class
eegdash.dataset.NM000111(cache_dir, query=None, s3_bucket=None, **kwargs)
Bases: EEGDashDataset
Author (year)
Canonical
Importable asNM000111
Sourceeegdash/dataset/registry.py · [source ↗]
class eegdash.dataset.NM000111(cache_dir: str, query: dict | None = None, s3_bucket: str | None = None, **kwargs)[source]#

ISRUC-Sleep: A comprehensive public dataset for sleep researchers.

Study:

nm000111 (NeMAR)

Author (year):

Canonical:

Also importable as: NM000111.

Modality: eeg; Subject type: Unknown. Subjects: 118; recordings: 126; tasks: 1.

Parameters:
  • cache_dir (str | Path) – Directory where data are cached locally.

  • query (dict | None) – Additional MongoDB-style filters to AND with the dataset selection. Must not contain the key dataset.

  • s3_bucket (str | None) – Base S3 bucket used to locate the data.

  • **kwargs (dict) – Additional keyword arguments forwarded to EEGDashDataset.

data_dir#

Local dataset cache directory (cache_dir / dataset_id).

Type:

Path

query#

Merged query with the dataset filter applied.

Type:

dict

records#

Metadata records used to build the dataset, if pre-fetched.

Type:

list[dict] | None

Notes

Each item is a recording; recording-level metadata are available via dataset.description. query supports MongoDB-style filters on fields in ALLOWED_QUERY_FIELDS and is combined with the dataset filter. Dataset-specific caveats are not provided in the summary metadata.

References

OpenNeuro dataset: https://openneuro.org/datasets/nm000111 NeMAR dataset: https://nemar.org/dataexplorer/detail?dataset_id=nm000111 DOI: https://doi.org/10.82901/nemar.nm000111

Examples

>>> from eegdash.dataset import NM000111
>>> dataset = NM000111(cache_dir="./data")
>>> recording = dataset[0]
>>> raw = recording.load()
__init__(cache_dir: str, query: dict | None = None, s3_bucket: str | None = None, **kwargs)[source]#
save(path: str, overwrite: bool = False, offset: int = 0)[source]#

Save datasets to files by creating one subdirectory for each dataset:

path/
    0/
        0-raw.fif | 0-epo.fif
        description.json
        raw_preproc_kwargs.json (if raws were preprocessed)
        window_kwargs.json (if this is a windowed dataset)
        window_preproc_kwargs.json  (if windows were preprocessed)
        target_name.json (if target_name is not None and dataset is raw)
    1/
        1-raw.fif | 1-epo.fif
        description.json
        raw_preproc_kwargs.json (if raws were preprocessed)
        window_kwargs.json (if this is a windowed dataset)
        window_preproc_kwargs.json  (if windows were preprocessed)
        target_name.json (if target_name is not None and dataset is raw)
Parameters:
  • path (str) –

    Directory in which subdirectories are created to store

    -raw.fif | -epo.fif and .json files to.

  • overwrite (bool) – Whether to delete old subdirectories that will be saved to in this call.

  • offset (int) – If provided, the integer is added to the id of the dataset in the concat. This is useful in the setting of very large datasets, where one dataset has to be processed and saved at a time to account for its original position.

Access modesMNE → braindecode → PyTorch → ML
.rawMNE Raw object — standard tools (filter, epoch, ICA, plot_psd).mne
DataLoaderWraps the windowed dataset into a PyTorch DataLoader; supports parallel workers and on-the-fly augmentations.pytorch
Zarr cacheOptional braindecode Zarr mirror for fast resume; persisted to cache_dir.zarr
Hugging FaceNo per-dataset mirror published yet — browse the EEGDash org listing for sibling datasets. See the datasets loader API.huggingface
Croissant 1.0Machine-readable JSON-LD descriptorNM000111.croissant.json (MLCommons schema, ingestible by PyTorch / TensorFlow / JAX).mlcommons
Examples using EEGDashcurated · start here

Swap any load_dataset(...) call for nm000111 to reproduce the tutorial on this dataset.

Citation

Sirvan Khalighi, Teresa Sousa, Jose Moutinho Santos, Urbano Nunes (2016). ISRUC-Sleep: A comprehensive public dataset for sleep researchers.. 10.82901/nemar.nm000111

Provenance

¹Contributed to nemar in BIDS format.

²Curated & ingested by the EEGDash catalog; see CITATION.cff for canonical reference.

³Persistent identifier: 10.82901/nemar.nm000111.

BIDS
BIDS 1.7.0
Sidecars
events · events.json · channels · eeg.json
Provenance
Machine-readable
Mirrors

See Also#