Research Article

Combining algorithm and data-level approaches to handle highly imbalanced soil data

DOI: 10.1080/02571862.2026.2652489
Author(s): Edward SmitNorth-West University, South Africa, George Munnik van ZijlNorth-West University, South Africa,

Abstract

The demand for detailed soil maps is increasing over time as they are considered a key element in sustainable development. However, imbalanced legacy soil data negatively affects digital soil mapping (DSM) accuracy, resulting in the loss of important minority classes. This study assessed different approaches using modern machine learning (ML) algorithms and upsampling techniques, testing four ML algorithms in their ability to predict two different soil classification systems prior to and after data balancing using four upsampling approaches (Synthetic Minority Oversampling Technique (SMOTE), Adaptive Synthetic Sampling Approach, Gaussian Noise and Random Oversampling) during model training. The results indicated that the algorithm approach is insufficient to handle highly imbalanced soil data; however, using balanced soil datasets in combination with ML algorithms led to a relative increase of 17% in predictive accuracy with kappa coefficient values exceeding 0.66 for both soil classification systems. The inherent taxonomical structure of soil classification system also affects soil class predictive accuracy. Care should be taken to account for catchment scales, distribution of data as well as spatial realism. It is concluded that a combined DSM approach using SMOTE and robust ML algorithms, such as Random Forest, can handle highly imbalanced soil datasets using different soil classification systems.

Get new issue alerts for South African Journal of Plant and Soil