Authors: Ambuj Kumar Misra
Abstract: Accurate crop yield prediction is fundamental to food security planning, resource optimization, and climate resilience policy. This study presents a comprehensive secondary data approach to predicting corn (Zea mays L.) yields across the contiguous United States by integrating multi-source datasets including National Oceanic and Atmospheric Administration (NOAA) climate records, the Soil Survey Geographic Database (SSURGO), USDA National Agricultural Statistics Service (USDA-NASS) historical yield data, and MODIS-derived Normalized Difference Vegetation Index (NDVI) values spanning 2000–2022. We evaluate and compare six predictive modeling frameworks—Linear Regression, Support Vector Regression (SVR), Random Forest (RF), Gradient Boosting (GB), Long Short-Term Memory (LSTM) neural networks, and a hybrid CNN-LSTM ensemble. The hybrid CNN-LSTM model achieved the highest predictive accuracy with an R² of 0.93 and a Root Mean Square Error (RMSE) of 5.4 bu/acre, substantially outperforming the baseline linear regression (R² = 0.61, RMSE = 18.4 bu/acre). Growing Degree Days, summer precipitation, and soil organic matter were identified as the three most influential predictors. Results demonstrate that rigorously curated secondary data, when combined with advanced machine learning architectures, can yield operationally reliable crop forecasts at county to regional scales without requiring expensive field campaigns. Implications for agricultural decision-making, early warning systems, and climate adaptation planning are discussed.