Public health risk stratification using hybrid machine learning: a reproducible analysis of performance, stability, and risk attribution
Keywords:
Artificial intelligence, Explainable risk attribution, Hybrid machine learning, Public health data integration, Risk stratificationAbstract
Risk stratification in public health involves organizing heterogeneous healthrelated signals into consistent representations that support population-level analysis. In large-scale datasets, such as National Health and Nutrition Examination Survey (NHANES) and Behavioral Risk Factor Surveillance System (BRFSS), the integration of clinical, biometric, behavioral, and self-reported variables introduces structural variability that challenges conventional modeling approaches. This study proposes a hybrid learning framework that combines linear and nonlinear components to analyze induced risk representations derived from multidimensional health data. The model is evaluated using NHANES 2017–2018, BRFSS 2019, and an Integrated Public Health Dataset constructed through semantic harmonization of both sources. The experimental design is based on a controlled formulation in which a continuous risk index is constructed from the available variables and discretized into ordinal classes using quantiles, enabling systematic analysis of how models approximate structured partitions of the input space rather than predicting independent clinical outcomes. The results show that the hybrid scheme maintains consistent macro F1 and macro-ROC-AUC values across all scenarios with low fold-to-fold variability, reflecting the regularity of the induced class structure rather than predictive generalization. Attribution analysis reveals that the organization of the risk representation varies according to the nature of the data, with concentrated patterns in clinical signals, distributed contributions in behavioral variables, and intermediate structures in the integrated dataset. These findings demonstrate that hybrid schemes provide a stable and interpretable framework for analyzing the structural organization of risk in heterogeneous public health data.