Using A Polynomial Regression Machine Learning Model To Predict Depression Severity Among People Living With HIV

Uncategorized

Authors: Geofrey Nyabuto, Peters Anselemo Ikoha, Samuel Mungai Mbuguah

Abstract: Background: Binary depression screening does not distinguish patients with mild symptoms from those with clinically urgent symptom burden. Among people living with HIV (PLHIV), depressive-symptom severity may reflect nonlinear interactions among anxiety, immune status, treatment adherence, stigma, behavioural exposures, and demographic characteristics. Routine electronic medical records (EMRs) provide an opportunity to model these relationships using computationally reproducible methods. Objective: To develop and internally validate a polynomial regression machine learning model for predicting continuous PHQ-9 depressive-symptom severity among PLHIV using routine HIV-care EMR data. Methods: A cross-sectional patient-level analytical dataset was constructed from 54,301 de-identified EMR records from Bungoma and Busia counties, Kenya. Fifteen clinical, treatment, psychosocial, behavioural, and demographic predictors were processed using training-derived imputation, encoding, log transformation, and standardisation. Stratified training (n=38,010), validation (n=8,145), and held-out test (n=8,146) partitions were used. Ordinary least-squares regression with degree-1, degree-2, and degree-3 polynomial feature expansions was compared using R², adjusted R², mean absolute error (MAE), mean squared error (MSE), and root mean squared error (RMSE). Results: The degree-3 model comprised 816 fitted parameters and achieved the strongest held-out performance: R² 0.753, adjusted R² 0.730, MAE 1.808, MSE 6.359, and RMSE 2.522 PHQ-9 points. The linear model achieved R² 0.729, MAE 1.939, and RMSE 2.641. Mean and median cubic-model residuals were 0.040 and 0.005, respectively, although the maximum positive residual was 15.186, indicating important underprediction in a small number of high-severity cases. The largest reported terms were anxiety × CD4 (β=−0.444) and anxiety² (β=0.413). Conclusions: Degree-3 polynomial regression modestly improved prediction of PHQ-9 depressive-symptom severity over linear and quadratic alternatives. Its average error may support broad risk stratification, but threshold crossing, model complexity, coefficient instability, and severe case underprediction preclude autonomous clinical use.

DOI: http://doi.org/

× How can I help you?