A Comparative Study: Machine Learning-Driven Yield Prediction for Turmeric Cultivation in Erode
Main Article Content
Abstract
In Erode, Tamil Nadu, which produces around 60% of India's total turmeric and is known as the "Turmeric Capital of the World," accurate forecasting of turmeric yield is crucial. In order to predict turmeric yield using multi-source agronomic data, this study offers a thorough comparative analysis of four cutting-edge machine learning models: Random Forest, Extreme Gradient Boosting, General Regression Neural Network and Long Short-Term Memory networks. The dataset includes meteorological variables such as rainfall, temperature and relative humidity and soil physicochemical parameters such as pH, nitrogen, phosphorus, potassium and moisture content, gathered over multiple crop seasons from three major taluks in the Erode district that grow turmeric. Z-score normalization, KNN-based methods for missing value imputation, and temporal lag feature engineering for sequential models were all part of the data preprocessing. Root Mean Square Error, Mean Absolute Error, and the Coefficient of Determination (R²) were used to assess the model's performance. Experimental results demonstrate that XGBoost outperforms Random Forest, GRNN and LSTM in terms of prediction accuracy, achieving lower error margins and a notably high coefficient of determination. The ensemble boosting technique, regularization procedures and intrinsic ability to represent intricate non-linear relationships between soil and climatic factors are responsible for XGBoost's exceptional performance. In order to help farmers, extension agents and agricultural officials make well-informed crop management decisions in the Erode turmeric belt, this research offers a scalable, data-driven methodology.