EEGdash›NeMAR›NM000287
Iss. 287 · 203 subjects · 540 recordings · CC-BY-NC-SA-4.0
Dataset Brief · Muse Sleep-Onset EEG — EEG/EMG Foundation Challenge 2026, Tra…

NM000287: eeg dataset, 203 subjects#

Muse Sleep-Onset EEG — EEG/EMG Foundation Challenge 2026, Track 03

Access recordings and metadata through EEGDash.

Citation: Muse Team (20). Muse Sleep-Onset EEG — EEG/EMG Foundation Challenge 2026, Track 03. 10.82901/nemar.nm000287

Modality: eeg Subjects: 203 Recordings: 540 License: CC-BY-NC-SA-4.0 Source: nemar

Metadata: Complete (100%)

203-participant EEG dataset — Muse Sleep-Onset EEG — EEG/EMG Foundation Challenge 2026, Track 03.

EEG · 4 ch128 HzBIDS 1.7.0Task · sleeponset5 sessions
Layer 01Study
What was asked
Hypothesis, independent & dependent variables, paradigm, cohort, and the editorial caveats around what the recordings can and cannot answer.
Layer 02Signal · BIDS
What was recorded
Sidecars, channels & electrodes, coordinate system, event semantics, and quality stats from the NEMAR pipeline when available.
Layer 03Training · ML
What you can train on
Recommended access modes — MNE Raw, braindecode windows, PyTorch DataLoader — plus the targets the metadata makes addressable.
§ 01Access · Get started

Quickstart#

Install

pip install eegdash

Access the data

from eegdash.dataset import NM000287

dataset = NM000287(cache_dir="./data")
# Get the raw object of the first recording
raw = dataset.datasets[0].raw
print(raw.info)

Filter by subject

dataset = NM000287(cache_dir="./data", subject="01")

Advanced query

dataset = NM000287(
    cache_dir="./data",
    query={"subject": {"$in": ["01", "02"]}},
)

Iterate recordings

for rec in dataset:
    print(rec.subject, rec.raw.info['sfreq'])

If you use this dataset in your research, please cite the original authors.

BibTeX

@dataset{nm000287,
  title = {Muse Sleep-Onset EEG — EEG/EMG Foundation Challenge 2026, Track 03},
  author = {Muse Team},
  doi = {10.82901/nemar.nm000287},
  url = {https://doi.org/10.82901/nemar.nm000287},
}
§ 02Study · The README

About This Dataset#

This dataset contains 540 at-home EEG recordings from 203 participants, collected

by Muse for Track 03 of the EEG/EMG Foundation Challenge 2026. The task is to predict the time remaining until the first N2 sleep epoch using only EEG recorded up to the time of prediction.

The recordings total approximately 157.52 hours and last about 6 to 30 minutes

each. All have four channels (TP9, AF7, AF8, TP10) sampled at 128 Hz, with one annotation marking the first N2 epoch. Full sleep-stage annotations are not included. Competition · Track 03 on Codabench

DOI

Muse sleep-onset EEG

Recordings and split

All 540 recordings are included in the NEMAR deposit, nm000287. The supplied split is stored in each participant’s sub-<label>_sessions.tsv:


View full README

DOI

Muse sleep-onset EEG

Recordings and split

All 540 recordings are included in the NEMAR deposit, nm000287. The supplied split is stored in each participant’s sub-<label>_sessions.tsv:

| Split | Recordings | Participants |
| --- | ---: | ---: |
| Train | 500 | 203 |
| Test | 40 | 40 |

Every test participant also appears in training. This split therefore evaluates new recordings from known participants, not generalization to unseen people.

It reproduces the supplied splits.csv; the assignments are included in the session tables, so that external file is not needed to use the deposit.

The competition evaluates both seen and unseen participants using weighted binned mean absolute error (W-bMAE). This deposit does not contain an unseen-person test cohort. Follow the competition code for evaluation windows, target clipping and scoring.

Read the data

Recordings use EEG-BIDS with BrainVision .vhdr, .vmrk and .eeg files.

Read them with MNE-BIDS, which applies the units and scaling from the header and loads the accompanying BIDS metadata:

from mne_bids import BIDSPath, read_raw_bids
path = BIDSPath(
    root="nm000287", subject="001", session="001", task="sleeponset",
    datatype="eeg", suffix="eeg", extension=".vhdr",

)
raw = read_raw_bids(path)

Set root to your local dataset directory. The binary EEG contains multiplexed 32-bit floats; reading it without the header scaling gives incorrect amplitudes.

Each events.tsv contains one n2_onset event (value=1). Its onset is in seconds from recording start, and sample is a zero-based sample index.

BrainVision marker positions are one-based. A duration of zero marks an instant, not the length of an N2 epoch. The marker’s Stimulus label is an export convention.

The N2 scoring method is not documented. Session tables also contain sample counts, durations, N2 onset and signal-quality flags; sessions.json defines these columns. participants.tsv contains recording counts and total durations. Demographics and acquisition dates are unavailable.

Session onset times are relative to recording start, not necessarily lights out. In EEG sidecars, RecordingDuration is the time to the last sample, (number_of_samples - 1) / 128; session durations use number_of_samples / 128.

Keep recording length out of the model

Every recording ends exactly 300 seconds after N2 onset. Knowing the full recording length therefore reveals the target without using EEG:

recording start                 first N2                 recording end
         |---------------------------|---------------------------|

      0                         onset                    onset + 300 s

At time t: use EEG up to t to predict onset - t.

For causal evaluation, hide total duration, sample counts, end-of-file information, N2 annotations and future EEG from the model. Whole-recording quality summaries and participant total durations must also remain outside the inputs. Preprocessing must not use future samples. These files do not enforce those restrictions; the evaluation pipeline must enforce them.

Acquisition and signal quality

The curator identified the hardware as Muse S family. The exact generation, firmware, reference and ground are unconfirmed. Headers identify pybv 0.7.5 as the export software, not the acquisition software. The stored rate is 128 Hz.

Downsampling from 256 Hz is a curator-supplied assumption; the original rate, resampling method and anti-aliasing filter have not been verified.

Prior filtering is unknown. HardwareFilters: n/a means that this information is unavailable. The supplied 60 Hz power-line setting and channel cutoff fields have been preserved but not independently verified. No filtering, resampling or signal correction was performed during metadata preparation. A screen of all 2,160 channel-recordings found no nonfinite samples or exactly constant aligned two-second windows. Median Welch spectra flagged 726 channel-recordings at 50 Hz and 958 at 60 Hz, using peaks more than 10 dB above adjacent bands. In 217 channel-recordings, more than 1% of samples exceeded 500 microvolts in absolute amplitude. At the recording level, 200 had 50 Hz flags, 294 had 60 Hz flags and 128 had amplitude flags; these groups overlap.

The flags identify recordings to inspect, not a diagnosis of artifacts. No isolated 50/60 Hz dip exceeded 10 dB below both spectral shoulders; apparent 60 Hz suppression against a combined baseline can reflect broad roll-off.

These spectra do not establish whether a notch filter was previously applied. All source channels are marked good, but those labels are not independent quality checks. Per-recording flags are described in SubjectArtefactDescription and the session tables. Detailed audit scripts and spectra are not included in this deposit.

Electrode and anatomical-landmark coordinates are identical across recordings, in metres in the CapTrak frame. Their provenance is unknown; they should not be treated as participant-specific measurements. Exact binary comparisons found no duplicate EEG files, including across splits. Partial overlap and transformed duplicates were not tested. Missing dates prevent checking chronological separation.

Credit, consent and validation

Credit Muse Team and cite dataset nm000287 with the version used. The data are licensed under CC-BY-NC-SA-4.0, matching the emg2pose release.

Muse determined internally that this collection was exempt from ethics review. The depositor confirmed authorization and participant consent for sharing, absence of identifiable personal information, and destruction of the re-identification key.

Before upload, BIDS validation passed with no errors. All 540 recordings were checked for consistent splits, channel order, sample counts and event timing.

Warnings remain for unavailable recommended metadata and collective authorship; no HED tags are used. To validate a local copy, run:

nemar dataset validate nm000287

For the file conventions, see EEG-BIDS (Pernet et al., 2019) and MNE-BIDS (Appelhoff et al., 2019).

NEMAR Metadata#

[![DOI](https://img.shields.io/badge/DOI-10.82901%2Fnemar.nm000287-blue)](https://doi.org/10.82901/nemar.nm000287) # Muse sleep-onset EEG This dataset contains 540 at-home EEG recordings from 203 participants, collected by Muse for Track 03 of the EEG/EMG Foundation Challenge 2026. The task is to predict the time remaining until the first N2 sleep epoch using only EEG recorded up to the time of prediction. The recordings total approximately 157.52 hours and last about 6 to 30 minutes each. All have four channels (TP9, AF7, AF8, TP10) sampled at 128 Hz, with one annotation marking the first N2 epoch. Full sleep-stage annotations are not included. [Competition](https://neural-interfaces26.github.io/tracks.html) · [Track 03 on Codabench](https://www.codabench.org/competitions/17983/) ## Recordings and split All 540 recordings are included in the NEMAR deposit, nm000287. The supplied split is stored in each participant’s sub-<label>_sessions.tsv: | Split | Recordings | Participants | | — | —: | —: | | Train | 500 | 203 | | Test | 40 | 40 | Every test participant also appears in training. This split therefore evaluates new recordings from known participants, not generalization to unseen people. It reproduces the supplied splits.csv; the assignments are included in the session tables, so that external file is not needed to use the deposit. The competition evaluates both seen and unseen participants using weighted binned mean absolute error (W-bMAE). This deposit does not contain an unseen-person test cohort. Follow the competition code for evaluation windows, target clipping and scoring. ## Read the data Recordings use EEG-BIDS with BrainVision .vhdr, .vmrk and .eeg files. Read them with MNE-BIDS, which applies the units and scaling from the header and loads the accompanying BIDS metadata: ```python from mne_bids import BIDSPath, read_raw_bids path = BIDSPath(

root=”nm000287”, subject=”001”, session=”001”, task=”sleeponset”, datatype=”eeg”, suffix=”eeg”, extension=”.vhdr”,

) raw = read_raw_bids(path) ``` Set root to your local dataset directory. The binary EEG contains multiplexed 32-bit floats; reading it without the header scaling gives incorrect amplitudes. Each events.tsv contains one n2_onset event (value=1). Its onset is in seconds from recording start, and sample is a zero-based sample index. BrainVision marker positions are one-based. A duration of zero marks an instant, not the length of an N2 epoch. The marker’s Stimulus label is an export convention. The N2 scoring method is not documented. Session tables also contain sample counts, durations, N2 onset and signal-quality flags; sessions.json defines these columns. participants.tsv contains recording counts and total durations. Demographics and acquisition dates are unavailable. Session onset times are relative to recording start, not necessarily lights out. In EEG sidecars, RecordingDuration is the time to the last sample, (number_of_samples - 1) / 128; session durations use number_of_samples / 128. ## Keep recording length out of the model Every recording ends exactly 300 seconds after N2 onset. Knowing the full recording length therefore reveals the target without using EEG: ```text recording start first N2 recording end

|---------------------------|—————————| 0 onset onset + 300 s

At time t: use EEG up to t to predict onset - t. ` For causal evaluation, hide total duration, sample counts, end-of-file information, N2 annotations and future EEG from the model. Whole-recording quality summaries and participant total durations must also remain outside the inputs. Preprocessing must not use future samples. These files do not enforce those restrictions; the evaluation pipeline must enforce them. ## Acquisition and signal quality The curator identified the hardware as Muse S family. The exact generation, firmware, reference and ground are unconfirmed. Headers identify pybv 0.7.5 as the export software, not the acquisition software. The stored rate is 128 Hz. Downsampling from 256 Hz is a curator-supplied assumption; the original rate, resampling method and anti-aliasing filter have not been verified. Prior filtering is unknown. `HardwareFilters: n/a` means that this information is unavailable. The supplied 60 Hz power-line setting and channel cutoff fields have been preserved but not independently verified. No filtering, resampling or signal correction was performed during metadata preparation. A screen of all 2,160 channel-recordings found no nonfinite samples or exactly constant aligned two-second windows. Median Welch spectra flagged 726 channel-recordings at 50 Hz and 958 at 60 Hz, using peaks more than 10 dB above adjacent bands. In 217 channel-recordings, more than 1% of samples exceeded 500 microvolts in absolute amplitude. At the recording level, 200 had 50 Hz flags, 294 had 60 Hz flags and 128 had amplitude flags; these groups overlap. The flags identify recordings to inspect, not a diagnosis of artifacts. No isolated 50/60 Hz dip exceeded 10 dB below both spectral shoulders; apparent 60 Hz suppression against a combined baseline can reflect broad roll-off. These spectra do not establish whether a notch filter was previously applied. All source channels are marked `good`, but those labels are not independent quality checks. Per-recording flags are described in `SubjectArtefactDescription` and the session tables. Detailed audit scripts and spectra are not included in this deposit. Electrode and anatomical-landmark coordinates are identical across recordings, in metres in the CapTrak frame. Their provenance is unknown; they should not be treated as participant-specific measurements. Exact binary comparisons found no duplicate EEG files, including across splits. Partial overlap and transformed duplicates were not tested. Missing dates prevent checking chronological separation. ## Credit, consent and validation Credit **Muse Team** and cite dataset `nm000287` with the version used. The data are licensed under [CC-BY-NC-SA-4.0](LICENSE), matching the emg2pose release. Muse determined internally that this collection was exempt from ethics review. The depositor confirmed authorization and participant consent for sharing, absence of identifiable personal information, and destruction of the re-identification key. Before upload, BIDS validation passed with no errors. All 540 recordings were checked for consistent splits, channel order, sample counts and event timing. Warnings remain for unavailable recommended metadata and collective authorship; no HED tags are used. To validate a local copy, run: ```sh nemar dataset validate nm000287 ` For the file conventions, see [EEG-BIDS (Pernet et al., 2019)](https://doi.org/10.1038/s41597-019-0104-8) and [MNE-BIDS (Appelhoff et al., 2019)](https://doi.org/10.21105/joss.01896).

License: CC-BY-NC-SA-4.0

Authors:

  • Muse Team

Versions:

Version

DOI

Released

current

10.82901/nemar.nm000287

§ 03Cohort · Participants

Cohort#

Dataset Statistics#

Channel counts: 4 ch (n=540 recordings)

Sampling frequencies: 128.0 Hz (n=540 recordings)

Total recording duration: 157 h

§ 04Signal · Electrodes & trace

Signal · Electrodes & live trace#

Fig. 01 Signal & montage 4 ch · EEG · 128 Hz · 203 subjects, 540 recordings
Live trace viewer — sub-001 · ses-001 · task-sleeponset

Showing one representative recording out of 203 subjects and 540 recordings in this dataset. Browse the full set on OpenNeuro; drop any other _eeg.{set,edf,bdf,vhdr} file onto the viewer (or pass ?eeg=<url>) to inspect it.

Electrode layout — EEG · 4 sensors — 4 channels

NEMAR Processing Statistics#

The plots below are generated by NEMAR’s automated EEG pipeline. The histogram shows pipeline success for data cleaning and ICA decomposition, the percentage of data frames and EEG channels retained after artefact removal, line noise per channel (RMS, dB), and the age/gender distribution of participants.

HED event descriptors word cloud HED event descriptors word cloud — NM000287
§ 05Manifest · BIDS tree

Manifest#

File Explorer#

Browse the BIDS file structure of this dataset. Records are fetched on demand from the EEGDash catalog the first time you open the explorer.

Recordings—
Files—
Subjects—
Modalities—
Click to load file structure…
Full dataset metadata table

Dataset ID

NM000287

Title

Muse Sleep-Onset EEG — EEG/EMG Foundation Challenge 2026, Track 03

Author (year)

—

Canonical

—

Importable as

NM000287

Year

20

Authors

Muse Team

License

CC-BY-NC-SA-4.0

Citation / DOI

10.82901/nemar.nm000287

Source links

OpenNeuro | NeMAR | Source URL

Copy-paste BibTeX
@dataset{nm000287,
  title = {Muse Sleep-Onset EEG — EEG/EMG Foundation Challenge 2026, Track 03},
  author = {Muse Team},
  doi = {10.82901/nemar.nm000287},
  url = {https://doi.org/10.82901/nemar.nm000287},
}
§ 06API · Programmatic access

API Reference#

Signature
eegdash.dataset
class
eegdash.dataset.NM000287(cache_dir, query=None, s3_bucket=None, **kwargs)
Bases: EEGDashDataset
Author (year)—
Canonical—
Importable asNM000287
Sourceeegdash/dataset/registry.py · [source ↗]
class eegdash.dataset.NM000287(cache_dir: str, query: dict | None = None, s3_bucket: str | None = None, **kwargs)[source]#

Muse Sleep-Onset EEG — EEG/EMG Foundation Challenge 2026, Track 03

Study:

nm000287 (NeMAR)

Author (year):

—

Canonical:

—

Also importable as: NM000287.

Modality: eeg; Subject type: Unknown. Subjects: 203; recordings: 540; tasks: 1.

Parameters:
  • cache_dir (str | Path) – Directory where data are cached locally.

  • query (dict | None) – Additional MongoDB-style filters to AND with the dataset selection. Must not contain the key dataset.

  • s3_bucket (str | None) – Base S3 bucket used to locate the data.

  • **kwargs (dict) – Additional keyword arguments forwarded to EEGDashDataset.

data_dir#

Local dataset cache directory (cache_dir / dataset_id).

Type:

Path

query#

Merged query with the dataset filter applied.

Type:

dict

records#

Metadata records used to build the dataset, if pre-fetched.

Type:

list[dict] | None

Notes

Each item is a recording; recording-level metadata are available via dataset.description. query supports MongoDB-style filters on fields in ALLOWED_QUERY_FIELDS and is combined with the dataset filter. Dataset-specific caveats are not provided in the summary metadata.

References

OpenNeuro dataset: https://openneuro.org/datasets/nm000287 NeMAR dataset: https://nemar.org/dataexplorer/detail?dataset_id=nm000287 DOI: https://doi.org/10.82901/nemar.nm000287

Examples

>>> from eegdash.dataset import NM000287
>>> dataset = NM000287(cache_dir="./data")
>>> recording = dataset[0]
>>> raw = recording.load()
__init__(cache_dir: str, query: dict | None = None, s3_bucket: str | None = None, **kwargs)[source]#
save(path: str, overwrite: bool = False, offset: int = 0)[source]#

Save datasets to files by creating one subdirectory for each dataset:

path/
    0/
        0-raw.fif | 0-epo.fif
        description.json
        raw_preproc_kwargs.json (if raws were preprocessed)
        window_kwargs.json (if this is a windowed dataset)
        window_preproc_kwargs.json  (if windows were preprocessed)
        target_name.json (if target_name is not None and dataset is raw)
    1/
        1-raw.fif | 1-epo.fif
        description.json
        raw_preproc_kwargs.json (if raws were preprocessed)
        window_kwargs.json (if this is a windowed dataset)
        window_preproc_kwargs.json  (if windows were preprocessed)
        target_name.json (if target_name is not None and dataset is raw)
Parameters:
  • path (str) –

    Directory in which subdirectories are created to store

    -raw.fif | -epo.fif and .json files to.

  • overwrite (bool) – Whether to delete old subdirectories that will be saved to in this call.

  • offset (int) – If provided, the integer is added to the id of the dataset in the concat. This is useful in the setting of very large datasets, where one dataset has to be processed and saved at a time to account for its original position.

Access modesMNE → braindecode → PyTorch → ML
.rawMNE Raw object — standard tools (filter, epoch, ICA, plot_psd).mne
DataLoaderWraps the windowed dataset into a PyTorch DataLoader; supports parallel workers and on-the-fly augmentations.pytorch
Zarr cacheOptional braindecode Zarr mirror for fast resume; persisted to cache_dir.zarr
Hugging FaceNo per-dataset mirror published yet — browse the EEGDash org listing for sibling datasets. See the datasets loader API.huggingface
Croissant 1.0Machine-readable JSON-LD descriptor — NM000287.croissant.json (MLCommons schema, ingestible by PyTorch / TensorFlow / JAX).mlcommons
Examples using EEGDashcurated · start here

Swap any load_dataset(...) call for nm000287 to reproduce the tutorial on this dataset.

Citation

Muse Team (20). Muse Sleep-Onset EEG — EEG/EMG Foundation Challenge 2026, Track 03. 10.82901/nemar.nm000287

Provenance

¹Contributed to nemar in BIDS format.

²Curated & ingested by the EEGDash catalog; see CITATION.cff for canonical reference.

³Persistent identifier: 10.82901/nemar.nm000287.

BIDS
BIDS 1.7.0
Sidecars
events · events.json · channels · eeg.json
Provenance
CC-BY-NC-SA-4.0 · 10.82901/nemar.nm000287
Machine-readable
Mirrors

See Also#