pyhealth.metrics.bootstrap#

Confidence intervals for evaluation metrics, as clinical reporting guidelines such as TRIPOD+AI expect. Samples from one patient are correlated, so pass groups=patient_ids to resample whole patients (a cluster bootstrap) rather than individual samples. To compare two models, use paired_bootstrap_diff: both are scored on identical resamples, so its interval is for their difference.

from pyhealth.metrics import bootstrap_ci, paired_bootstrap_diff

ci = bootstrap_ci(y_true, y_prob, "pr_auc", groups=patient_ids, n_boot=1000)
# {'estimate': 0.41, 'lower': 0.35, 'upper': 0.47, 'n_boot': 1000, 'n_skipped': 0}

diff = paired_bootstrap_diff(y_true, y_prob_new, y_prob_baseline, "pr_auc",
                             groups=patient_ids)
# an interval that excludes 0 indicates a difference at the 5% level

metric is any name accepted by binary_metrics_fn, including the calibration metrics brier, oe_ratio, calibration_slope and calibration_intercept, or a callable metric(y_true, y_prob) -> float. Resamples that contain a single class are skipped and counted in n_skipped. See examples/calibration_and_bootstrap_ci.py.

pyhealth.metrics.bootstrap.bootstrap_ci(y_true, y_prob, metric, groups=None, n_boot=1000, seed=0, alpha=0.05)[source]#

Percentile bootstrap confidence interval for a binary metric.

Parameters:
  • y_true (ndarray) – True binary labels, shape (n,).

  • y_prob (ndarray) – Predicted probabilities, shape (n,).

  • metric (str | Callable[ndarray, ndarray, float]) – A name accepted by binary_metrics_fn (e.g. "roc_auc", "brier") or a callable metric(y_true, y_prob) -> float.

  • groups (Optional[ndarray]) – Optional cluster ids, shape (n,), e.g. patient ids. Whole groups are resampled with replacement, keeping all their samples.

  • n_boot (int) – Number of resamples drawn.

  • seed (int) – Random seed; results are deterministic given it.

  • alpha (float) – 1 - confidence level (0.05 gives a 95% interval).

Return type:

dict

Returns:

dict with estimate (metric on all data), lower and upper (percentile bounds), n_boot (resamples used) and n_skipped (resamples with a single class, which are skipped).

Examples

>>> import numpy as np
>>> from pyhealth.metrics.bootstrap import bootstrap_ci
>>> y_true = np.array([0, 0, 1, 1, 0, 1, 0, 1])
>>> y_prob = np.array([0.1, 0.3, 0.7, 0.8, 0.4, 0.6, 0.2, 0.9])
>>> patients = np.array([1, 1, 2, 2, 3, 3, 4, 4])
>>> ci = bootstrap_ci(y_true, y_prob, "brier", groups=patients, n_boot=200)
>>> sorted(ci)
['estimate', 'lower', 'n_boot', 'n_skipped', 'upper']
pyhealth.metrics.bootstrap.paired_bootstrap_diff(y_true, y_prob_a, y_prob_b, metric, groups=None, n_boot=1000, seed=0, alpha=0.05)[source]#

Bootstrap interval for metric(model A) - metric(model B) on the same data.

Both models are scored on identical resamples, so the interval reflects their difference rather than two independent uncertainties. An interval that excludes 0 indicates a difference at level alpha.

Parameters:
  • y_true (ndarray) – True binary labels, shape (n,).

  • y_prob_a (ndarray) – Model A’s predicted probabilities, shape (n,).

  • y_prob_b (ndarray) – Model B’s predicted probabilities, shape (n,).

  • metric (str | Callable[ndarray, ndarray, float]) – A binary_metrics_fn name or a callable, as in bootstrap_ci().

  • groups (Optional[ndarray]) – Optional cluster ids (e.g. patient ids) for a cluster bootstrap.

  • n_boot (int) – Number of resamples drawn.

  • seed (int) – Random seed; results are deterministic given it.

  • alpha (float) – 1 - confidence level.

Return type:

dict

Returns:

dict with estimate (A - B on all data), lower, upper, n_boot and n_skipped, as in bootstrap_ci().

Examples

>>> import numpy as np
>>> from pyhealth.metrics.bootstrap import paired_bootstrap_diff
>>> y_true = np.array([0, 0, 1, 1, 0, 1, 0, 1])
>>> model_a = np.array([0.1, 0.3, 0.7, 0.8, 0.4, 0.6, 0.2, 0.9])
>>> model_b = np.array([0.4, 0.5, 0.5, 0.6, 0.5, 0.4, 0.3, 0.7])
>>> diff = paired_bootstrap_diff(y_true, model_a, model_b, "roc_auc", n_boot=200)
>>> diff["estimate"] > 0
True