Authors: Professor Priyanka Mahajan, Rohit Balaji Pingale, Vaishnavi Laxman Shinde
Abstract: A India's electronic waste (e-waste) generation has increased rapidly, while formal collection and recycling infrastructure remains concentrated in major metropolitan areas. Tier-2 and Tier-3 cities continue to face challenges in planning efficient collection networks due to limited data-driven decision support. This paper presents a machine learning-based framework for estimating ward-level e-waste generation using publicly available demographic, economic, and administrative indicators. A Random Forest regressor is proposed as the primary prediction model and is benchmarked against Gradient Boosting and Ordinary Least Squares regression to evaluate predictive performance. The framework incorporates a reproducible data-engineering pipeline integrating data from the Central Pollution Control Board (CPCB), State Pollution Control Boards (SPCBs), Census records, and Urban Local Bodies (ULBs). To validate the proposed workflow before large-scale field deployment, a clearly labelled synthetic dataset is developed for model training and evaluation. Feature importance analysis is further employed to identify the key factors influencing e-waste generation, enabling practical insights for municipal authorities and Extended Producer Responsibility (EPR) stakeholders. The proposed approach offers a scalable, low-cost, and replicable methodology for supporting sustainable e-waste management, optimizing collection planning, and reducing the infrastructure gap between metropolitan and emerging urban regions in India.