pyhealth.datasets.EEGBCIDataset#

class pyhealth.datasets.EEGBCIDataset(root, dataset_name=None, config_path=None, subjects=None, runs=None, download=False, **kwargs)[source]#

Bases: BaseDataset

PhysioNet EEG Motor Movement/Imagery metadata dataset.

The source dataset is PhysioNet’s EEG Motor Movement/Imagery Dataset (eegmmidb), version 1.0.0, licensed under the Open Data Commons Attribution License v1.0. Cite Schalk (2009), https://doi.org/10.13026/C28G6P.

Parameters:
  • root (str) – Directory containing or receiving EEGBCI EDF files and metadata.

  • dataset_name (Optional[str]) – Optional dataset name prefix. Defaults to "eegbci".

  • config_path (Optional[str]) – Optional dataset configuration path.

  • subjects (Optional[list[int]]) – Subject identifiers to include. Defaults to [1, 2, 3].

  • runs (Optional[list[int]]) – Run identifiers to include. Defaults to runs 3 through 14.

  • download (bool) – Whether MNE may download missing EDF files.

  • **kwargs – Additional arguments forwarded to BaseDataset.

Raises:

FileNotFoundError – If a requested EDF is unavailable and downloading is disabled.

Examples

>>> dataset = EEGBCIDataset(
...     root="/path/to/eegbci", subjects=[1], runs=[3], download=True
... )
>>> dataset.stats()
prepare_metadata()[source]#

Reuse valid metadata or write rows for every requested EDF.

Raises:

FileNotFoundError – If a requested EDF is unavailable and downloading is disabled.

Return type:

None

property default_task: EEGMotorImageryEEGBCI#

Return the canonical supervised EEGBCI task.

Return type:

EEGMotorImageryEEGBCI

Returns:

An EEGMotorImageryEEGBCI task.

clean_tmpdir()#

Cleans up the temporary directory within the cache.

Return type:

None

create_tmpdir()#

Creates and returns a new temporary directory within the cache.

Returns:

The path to the new temporary directory.

Return type:

Path

get_patient(patient_id)#

Retrieves a Patient object for the given patient ID.

Parameters:

patient_id (str) – The ID of the patient to retrieve.

Returns:

The Patient object for the given ID.

Return type:

Patient

Raises:

AssertionError – If the patient ID is not found in the dataset.

property global_event_df: LazyFrame#

Returns the path to the cached event dataframe.

Returns:

The path to the cached event dataframe.

Return type:

Path

iter_patients(df=None)#

Yields Patient objects for each unique patient in the dataset.

Yields:

Iterator[Patient] – An iterator over Patient objects.

Return type:

Iterator[Patient]

load_data()#

Loads data from the specified tables.

Returns:

A concatenated lazy frame of all tables.

Return type:

dd.DataFrame

load_table(table_name)#

Loads a table and processes joins if specified.

Parameters:

table_name (str) – The name of the table to load.

Returns:

The processed Dask dataframe for the table.

Return type:

dd.DataFrame

Raises:
  • ValueError – If the table is not found in the config.

  • FileNotFoundError – If the source file (CSV/TSV or Parquet) for the table or join is not found.

set_task(task=None, num_workers=None, input_processors=None, output_processors=None, split=None)#

Processes the base dataset to generate the task-specific sample dataset. The cache structure is as follows:

{task_name}_{task_uuid}/        # Cached data for specific task based on task name, schema, and args
    task_df.ld/                 # Intermediate task dataframe based on schema
    samples_{proc_uuid}.ld/     # Final processed samples after applying processors
        schema.pkl              # Saved SampleBuilder schema
        split.npz               # Sample indices per split part (with split= only)
        *.bin                   # Processed sample files
Parameters:
  • task (Optional[BaseTask]) – The task to set. Uses default task if None.

  • num_workers (int) – Number of workers for multi-threading. Default is self.num_workers.

  • input_processors (Optional[Dict[str, FeatureProcessor]]) – Pre-fitted input processors. If provided, these will be used instead of creating new ones from task’s input_schema. Defaults to None.

  • output_processors (Optional[Dict[str, FeatureProcessor]]) – Pre-fitted output processors. If provided, these will be used instead of creating new ones from task’s output_schema. Defaults to None.

  • split (Optional[Split]) – Split the samples, e.g. by patient with PatientSplit, and fit every processor on the first (training) part only, so the other parts never shape preprocessing. The parts are used as the split returns them. Samples are streamed; nothing is loaded into memory beyond the per-sample index that processing already keeps. Defaults to None: fit on all samples and return one dataset, as before.

Returns:

The generated sample dataset, or, with split, a tuple with one dataset per part, training part first, whose processors were fitted on the training part.

Return type:

SampleDataset

Examples

>>> from pyhealth.datasets import PatientSplit
>>> train, val, test = dataset.set_task(  
...     task, split=PatientSplit(ratios=(0.7, 0.1, 0.2), seed=42)
... )
Raises:

AssertionError – If no default task is found and task is None.

stats()#

Prints statistics about the dataset.

Return type:

None

property unique_patient_ids: List[str]#

Returns a list of unique patient IDs.

Returns:

List of unique patient IDs.

Return type:

List[str]