Sunday A. Adetunji, Rhoda O. Oyewusi
Abstract
Preterm birth remains a major cause of neonatal morbidity and mortality worldwide. Electrohysterography (EHG), a noninvasive measure of uterine myoelectrical activity, has been studied for preterm-birth prediction, but performance estimates may be biased when segments from the same maternal record are split across training and validation data. We formalized the distinction between segment-level and patient-independent record-grouped validation, established a patient-independent benchmark on the Term-Preterm Electrohysterogram Database, and evaluated class-conditional conformal selective prediction. All 300 records (38 preterm) were analyzed across three prespecified regimes. A 92-feature elastic-net logistic model was evaluated with record-grouped nested cross-validation; preprocessing, model fitting, Platt calibration, and conformal estimation were confined to training data, and performance was calculated from strictly out-of-fold record-level predictions with 1,000 bootstrap resamples. Under patient-independent evaluation, AUROC was 0.493 (95% CI 0.467-0.520), AUPRC 0.122 (0.095-0.152), and Brier score 0.115 (0.094-0.137). AUROC was 0.514 at or before 26 weeks and 0.469 thereafter. At miscoverage alpha=0.10, marginal coverage was 0.897, abstention 72.7%, and singleton-prediction accuracy 0.624. These findings establish a record-separated reference benchmark for EHG prediction and a broader validation principle for segmented physiological data: the unit of resampling should correspond to the unit at which predictive performance is intended to generalize. Conformal prediction further quantifies when the available information supports a singleton classification and when uncertainty warrants deferral.