pyhealth.metrics.bootstrap#
Confidence intervals for evaluation metrics, as clinical reporting guidelines
such as TRIPOD+AI expect. Samples from one patient are correlated, so pass
groups=patient_ids to resample whole patients (a cluster bootstrap) rather
than individual samples. To compare two models, use paired_bootstrap_diff:
both are scored on identical resamples, so its interval is for their difference.
from pyhealth.metrics import bootstrap_ci, paired_bootstrap_diff
ci = bootstrap_ci(y_true, y_prob, "pr_auc", groups=patient_ids, n_boot=1000)
# {'estimate': 0.41, 'lower': 0.35, 'upper': 0.47, 'n_boot': 1000, 'n_skipped': 0}
diff = paired_bootstrap_diff(y_true, y_prob_new, y_prob_baseline, "pr_auc",
groups=patient_ids)
# an interval that excludes 0 indicates a difference at the 5% level
metric is any name accepted by binary_metrics_fn, including the
calibration metrics brier, oe_ratio, calibration_slope and
calibration_intercept, or a callable metric(y_true, y_prob) -> float.
Resamples that contain a single class are skipped and counted in n_skipped.
See examples/calibration_and_bootstrap_ci.py.
- pyhealth.metrics.bootstrap.bootstrap_ci(y_true, y_prob, metric, groups=None, n_boot=1000, seed=0, alpha=0.05)[source]#
Percentile bootstrap confidence interval for a binary metric.
- Parameters:
y_true (
ndarray) – True binary labels, shape (n,).y_prob (
ndarray) – Predicted probabilities, shape (n,).metric (
str|Callable[ndarray,ndarray,float]) – A name accepted bybinary_metrics_fn(e.g."roc_auc","brier") or a callablemetric(y_true, y_prob) -> float.groups (
Optional[ndarray]) – Optional cluster ids, shape (n,), e.g. patient ids. Whole groups are resampled with replacement, keeping all their samples.n_boot (
int) – Number of resamples drawn.seed (
int) – Random seed; results are deterministic given it.alpha (
float) – 1 - confidence level (0.05 gives a 95% interval).
- Return type:
- Returns:
dict with
estimate(metric on all data),lowerandupper(percentile bounds),n_boot(resamples used) andn_skipped(resamples with a single class, which are skipped).
Examples
>>> import numpy as np >>> from pyhealth.metrics.bootstrap import bootstrap_ci >>> y_true = np.array([0, 0, 1, 1, 0, 1, 0, 1]) >>> y_prob = np.array([0.1, 0.3, 0.7, 0.8, 0.4, 0.6, 0.2, 0.9]) >>> patients = np.array([1, 1, 2, 2, 3, 3, 4, 4]) >>> ci = bootstrap_ci(y_true, y_prob, "brier", groups=patients, n_boot=200) >>> sorted(ci) ['estimate', 'lower', 'n_boot', 'n_skipped', 'upper']
- pyhealth.metrics.bootstrap.paired_bootstrap_diff(y_true, y_prob_a, y_prob_b, metric, groups=None, n_boot=1000, seed=0, alpha=0.05)[source]#
Bootstrap interval for metric(model A) - metric(model B) on the same data.
Both models are scored on identical resamples, so the interval reflects their difference rather than two independent uncertainties. An interval that excludes 0 indicates a difference at level
alpha.- Parameters:
y_true (
ndarray) – True binary labels, shape (n,).y_prob_a (
ndarray) – Model A’s predicted probabilities, shape (n,).y_prob_b (
ndarray) – Model B’s predicted probabilities, shape (n,).metric (
str|Callable[ndarray,ndarray,float]) – Abinary_metrics_fnname or a callable, as inbootstrap_ci().groups (
Optional[ndarray]) – Optional cluster ids (e.g. patient ids) for a cluster bootstrap.n_boot (
int) – Number of resamples drawn.seed (
int) – Random seed; results are deterministic given it.alpha (
float) – 1 - confidence level.
- Return type:
- Returns:
dict with
estimate(A - B on all data),lower,upper,n_bootandn_skipped, as inbootstrap_ci().
Examples
>>> import numpy as np >>> from pyhealth.metrics.bootstrap import paired_bootstrap_diff >>> y_true = np.array([0, 0, 1, 1, 0, 1, 0, 1]) >>> model_a = np.array([0.1, 0.3, 0.7, 0.8, 0.4, 0.6, 0.2, 0.9]) >>> model_b = np.array([0.4, 0.5, 0.5, 0.6, 0.5, 0.4, 0.3, 0.7]) >>> diff = paired_bootstrap_diff(y_true, model_a, model_b, "roc_auc", n_boot=200) >>> diff["estimate"] > 0 True