Open-access Prediction of soil fertility properties in Southern Brazil via proximal sensing

Abstract

Although proximal sensing coupled with machine learning (ML) algorithms have been successful for characterizing soils, questions remain regarding their effectiveness under varied soil conditions. This study evaluated for the first time the efficiency of a portable X-ray fluorescence spectrometer (pXRF) to predict 17 soil fertility properties in Rio Grande do Sul (RS) state, Brazil, through ML algorithms. A total of 468 surface soil samples were analyzed by pXRF and by conventional (reference) methods. Six algorithms were employed: Projection Pursuit Regression, Partial Least Squares, Random Forest, Support Vector Machine, Extreme Gradient Boosting, and Cubist. Predictions accuracy was assessed using the coefficient of determination (R²), root mean square error, normalized root mean square error, residual prediction deviation (RPD) and Ratio of Performance to Interquartile Distance. Cubist and Random Forest outperformed other algorithms, reaching the following R² values: available/exchangeable Al (R² = 0.70), Ca (0.57), Mg (0.75), Mn (0.84), S (0.60), Cu (0.81), K (0.82), P (0.54), besides P-rem (0.80), H+Al (0.73), and total N (0.52). Predictions for organic carbon and available B, Fe, Na, Zn require further investigations. The pXRF combined with ML algorithms can accelerate decisions for agricultural management in RS state, Brazil, by optimizing soil analysis for improved crop management.

Key words
Portable X-ray fluorescence; soil properties; pedometrics; soil variation

INTRODUCTION

Efforts to increase food production must be intensified to meet the FAO projections for 2030. It is essential to focus on enhancing the productivity of existing agricultural areas, avoiding the conversion of areas under native vegetation into production systems, and promoting soil conservation (FAO et al. 2023). In this context, detailed information on soil properties becomes vital for the global development agenda, supporting more specific and sustainable management practices to optimize production (Benedet et al. 2021a). However, the lack of detailed soil data undermines the necessary support to effectively manage soils, optimizing production while ensuring the sustainable use of soil and water resources (Liu et al. 2022).

Brazil plays a crucial role in the global agricultural and forestry production systems, ranking among the largest producers of soybeans, corn, rice, coffee, ethanol, sugar, cellulose, and other commodities (Tanure et al. 2024), under contrasting soil conditions. The state of Rio Grande do Sul (RS), in particular, exhibits distinctive characteristics across its various physiographic regions. In the RS plateau, soils are deep and highly weathered; in the campaign region, soils are highly fertile; whereas in the coastal region, soils exhibit pronounced hydromorphic traits (Streck et al. 2018). To address productive challenges, RS seeks new sources of information to promote sustainable development. This includes alternatives to accelerate the characterization of diverse soil types (Moura-Bueno et al. 2019), expansion of pedological maps through similar physiographic areas (Bagatini et al. 2016, Pelegrino et al. 2016), mapping of organic carbon and soil carbon stocks (Bonfatti et al. 2016, Tornquist et al. 2024), and the use of spectral libraries for soil properties prediction (Moura-Bueno et al. 2020).

The demand for soil information to manage soil fertility of productive areas represents a continuous effort (Ribeiro et al. 2024). Soil organic carbon (OC), macronutrients (Ca, K, Mg, N, P, and S), micronutrients (B, Cu, Fe, Mn, and Zn), acidity (pH), potential acidity (H+Al), and aluminum (Al) require substantial effort when analyzed using conventional methods. In this gap, there is potential to optimize these processes through the use of proximal sensing, which enable rapid soil characterization, generating data that can be correlated with various soil properties (Silva et al. 2021), which still needs further investigations in different soil conditions. Under this proximal sensing approach, it has been possible to obtain soil properties data (e.g., soil fertility properties) more rapidly, cost effectively and without generation of chemical waste, accelerating decision makings on soil management practices, contributing to increase yields (Tavares et al. 2020b).

Among these sensors, portable X-ray fluorescence (pXRF) stands out for quantifying chemical elements in samples rapidly and without generating waste (Pelegrino et al. 2022). This process is based on the reaction of each element to energy stimulation across different ranges of the X-ray electromagnetic spectrum (Weindorf et al. 2014). The integration of these results with soil properties using machine learning techniques has enabled accurate predictions of soil texture (sand, silt, and clay content) (Silva et al. 2020, Tavares et al. 2020b, Weindorf & Chakraborty 2020) and fertility (Andrade et al. 2020b, 2023, Mancini et al. 2024, Silva et al. 2020, Zhu et al. 2011). Although some studies on this regard have used linear regressions to predict soil attributes, global research on machine learning estimating soil attributes using proximal sensing have outperformed linear regressions (Benedet et al. 2021a, Karami et al. 2024b, Wan et al. 2020, Mozaffari et al. 2024a), and highlighted the need for region-specific models (Brungard et al. 2021), which have not been developed for soils of Rio Grande do Sul state, southern Brazil.

This study employed six supervised learning algorithms commonly used in agricultural and environmental modeling. Partial Least Squares (PLS) and Projection Pursuit Regression (PPR) are linear techniques that reduce data dimensionality and are useful when variables are highly collinear (Wold et al. 2001, Friedman & Stuetzle 1981). Random Forest (RF) and Extreme Gradient Boosting (XGBoost) are ensemble tree-based methods that capture complex non-linear patterns and perform well with high-dimensional data (Breiman 2001, Chen & Guestrin 2016). Support Vector Machine (SVM) is effective for regression tasks involving non-linear relationships (Smola & Schölkopf 2004), while Cubist combines rule-based models with linear regression and has shown strong performance in predicting soil properties (Padarian et al. 2017).

In light of the above, this study aims to utilize pXRF to generate a regional dataset capable of training and validating machine learning models for predicting 17 soil fertility properties, as well as identifying the most relevant elements for predictions under these conditions. The soils of Rio Grande do Sul exhibit distinct attributes compared to other regions in Brazil due to its diverse geology, topography, and climate, which directly influence soil formation factors, rising questions on the efficiency of this approach for such soils. Thus, this study seeks not only to validate the use of pXRF in RS but also to provide a solid foundation for expanding the results obtained with this technology and integrating them into the national context. By investigating its effectiveness in soils with distinct environmental conditions and formation factors, such as those in RS, this work intends to advance the use of pXRF as a robust and precise tool for predicting soil fertility attributes, paving the way for its wider application in diverse areas of Brazil.

MATERIALS AND METHODS

Study area

Soil samples were collected in the state of Rio Grande do Sul (RS), located in southern Brazil (Figure 1). The region predominantly exhibits a humid subtropical oceanic climate with hot summers and no dry season (Cfa), transitioning to a temperate summer climate (Cfb) in higher altitudes (Alvares et al. 2013b). Annual precipitation ranges from 1,600 to 1,900 mm, with altitudes varying from 0 to 400 meters, and mean annual temperatures between 14°C and 17°C (Alvares et al. 2013a, Lima et al. 2023). Geologically, the study area covers two geomorphological provinces: the Peripheral Depression, part of the extensive Paraná Basin and composed predominantly of sandstones, siltstones, and mudstones; and the Coastal Plain, located at lower elevations (<40 m), characterized by unconsolidated sediments influenced by marine transgression and regression events.

Figure 1
Location map. (a) Relative position of Rio Grande do Sul (RS) in Brazil and South America. (b) Distribution of collected samples in RS. (c) Detail of altitude variation relative to collected samples.

Sampling and Laboratory Analyses

A total of 468 surface soil samples (0–20 cm) were collected across multiple regions of Rio Grande do Sul (RS), Brazil, in partnership with CMPC Celulose Rio Grandense. The depth of 0–20 cm corresponds to the most relevant soil layer adopted for evaluating soil fertility and management responses, especially in systems involving lime and fertilizer application, as it is typically affected by agricultural practices, root activity, and nutrient cycling. Field prospection followed the free walking transect method under eucalyptus plantations, where researchers observed changes in soil class and properties in situ. Soil sampling points were selected whenever visual or morphological variations in soil attributes were noted, ensuring high representativeness across contrasting soil types.

The samples were air-dried and sieved to a 2 mm fraction for chemical analyses. Soil pH in water was measured at a 1:2.5 ratio following the protocol recommended by Teixeira et al. (2017)Remaining phosphorus (P-rem) was determined following the method of Alvarez et al. (2000), and potential acidity (H+Al³⁺) was assessed using calcium acetate extraction followed by titration with NaOH (Vettori 1969). Organic carbon was determined using wet oxidation with potassium dichromate in a sulfuric acid medium (Walkley & Black 1934). Exchangeable Ca²⁺, Mg²⁺, and Al³⁺ were quantified as per McLean et al. (1958), while available K⁺ was extracted using the Mehlich-1 method (Mehlich 1953). Exchangeable Na⁺ was extracted with 0.05 mol L⁻¹ HCl.

All samples were analyzed using a pXRF Tracer 5g (Bruker Analytical Instrumentation, Billerica, MA, USA). Scans were performed in triplicate in Soil mode for 60 seconds each. The procedures followed the methodology described by Silva et al. (2021) and Weindorf & Chakraborty, (2020). To assess the analytical accuracy of the pXRF results, three certified reference materials (CRMs) were analyzed: NIST 2710a and NIST 2711a, provided by the National Institute of Standards and Technology, and a check sample (CS) supplied by the equipment manufacturer, in order to calculate the recovery values per element detected. For that purpose, these CRM samples were scanned by pXRF using the same settings to be applied to the soil samples. Then, the content of each element detected by pXRF was divided by the certified content of that element in each CRM sample to calculate the recovery values of each element detected by pXRF (Recovery value (%) = 100 x content of element detected by pXRF/certified content of the same element). Recovery values are important to assess the accuracy of the equipment per element and also to inform readers about the possible over- or underestimation of contents determined by pXRF. The recovery values for the 13 elements detected by pXRF in all samples and, thus, used in this study for soil fertility predictions, are presented here following the order of recovery values obtained for NIST2710a, NIST2711a, and CS, and were: Al - 82/75/92; Ca – 36/40/–; Cr - –/101/–; Cu – 81/73/94; Fe – 73/67/85; K – 62/48/84; Mg (–/66/–); Mn – 67/60/81; P (210/224/–); Si – 58/49/85; Ti – 77/66/–; V – 50/29/–; Zn – 87/85/–; Zr (86/–/–). Some elements were not certified in all reference materials, and thus recovery values were not calculated in those cases. Despite some variation, especially for elements like P and Ca, the overall recovery rates for key elements used in modeling (such as Al, Cu, Fe, Mn, Zn, and Si) demonstrate a satisfactory degree of analytical reliability.

In total, 17 soil fertility properties were estimated using predictive models based on the 13 elements detected by pXRF in all samples (Al, Ca, Fe, K, Mg, P, Rb, Si, Sr, Ti, V, Zn, and Zr): pH, exchangeable Al, Ca, Mg, available P, K, S, B, Cu, Fe, Mn, Na, Zn, total N, remaining phosphorus (P-rem), potential acidity (H+Al), and organic carbon (OC). These properties were predicted using 13 elemental covariates obtained from pXRF analysis: Al, Ca, Fe, K, Mg, P, Rb, Si, Sr, Ti, V, Zn, and Zr. To evaluate the distributional behavior of the dataset, a Shapiro–Wilk normality test was performed for all 17 soil fertility attributes. The test was applied to assess whether each variable followed a Gaussian distribution, as this can influence model assumptions and preprocessing strategies. The results indicated that none of the variables followed a normal distribution (all p-values < 0.001), suggesting the presence of significant deviations from normality, likely due to the natural heterogeneity of soil chemical attributes.

Data Analysis and Modeling

The 17 soil fertility attributes (pH, exchangeable Al, Ca, Mg, available P, K, S, B, Cu, Fe, Mn, Na, Zn, total N, remaining phosphorus (P-rem), potential acidity (H+Al), and organic carbon (OC)) were estimated by the 13 elements detected by pXRF in all samples, using six machine learning algorithms: Projection Pursuit Regression (PPR) (Friedman & Stuetzle 1981), Partial Least Squares (PLS) (Mozaffari et al. 2024b), Random Forest (RF) (Boehmke & Greenwell 2019), Support Vector Machine (SVM) (Bennett & Campbell 2000), Extreme Gradient Boosting (XGB) (Friedman 2001), and Cubist (Wang & Witten 1997). For model training, 70% (328) of the dataset randomly separated were used, while the remaining 30% (140) were reserved for validation (Mozaffari et al. 2024a).

Regarding the six machine learning algorithms used to model soil fertility attributes based on elemental data from pXRF, Projection Pursuit Regression (PPR) is a nonparametric method that projects multidimensional data into lower-dimensional spaces and fits smooth functions along those projections to capture nonlinear patterns (Friedman & Stuetzle 1981). Partial Least Squares (PLS) is a linear regression technique that reduces dimensionality by extracting latent variables (components) that maximize covariance between predictors and the response variable (Wold et al. 2001). Random Forest (RF) is an ensemble of decision trees where each tree is built on a bootstrap sample and splits are chosen based on a random subset of features, improving prediction robustness and reducing overfitting (Breiman 2001). Support Vector Machine (SVM) regression maps data into high-dimensional space using kernel functions and finds the hyperplane that minimizes error within a defined margin (Smola & Schölkopf 2004). Extreme Gradient Boosting (XGB) is a sequential ensemble method that builds decision trees iteratively, where each tree corrects the residuals of the previous one, with regularization to prevent overfitting (Chen & Guestrin 2016). Lastly, Cubist combines rule-based models and linear regression by partitioning the predictor space using decision rules and fitting a linear model in each partition, offering good balance between interpretability and performance (Wang & Witten 1997, Kuhn & Quinlan 2011).

All predictive models were implemented in the R software environment (version 4.3.1) using the caret package (Kuhn 2008) as the primary modeling framework, which streamlines training, validation, and hyperparameter tuning of multiple algorithms. The following additional R packages were used to support specific methods: pls for Partial Least Squares (PLS) (Liland et al. 2024), randomForest for Random Forest (Breiman et al. 2002), e1071 for Support Vector Machines (SVM) (Meyer et al. 2024), XGBoost for Extreme Gradient Boosting (XGB) (Chen et al. 2025), and Cubist (Kuhn & Quinlan 2011) for rule-based regression modeling. For Projection Pursuit Regression (PPR), the native implementation within caret was used. All models were trained using 10-fold cross-validation and tuned using grid search to optimize performance based on RMSE and R² values.

To evaluate model performance, the root mean square error (RMSE) was used as a primary metric to assess the average magnitude of prediction errors. Since RMSE is scale-dependent, its interpretation was performed in conjunction with the coefficient of determination (R²) and the residual prediction deviation (RPD). Lower RMSE values generally indicate better performance, but they are most meaningful when compared across models predicting the same variable (Bellon-Maurel et al. 2010).

Model external validation was performed by using the 30% of samples not adopted for model training and calculating the coefficient of determination (R²) (Eq. 1), root mean square error (RMSE) (Eq. 2), residual prediction deviation (RPD) (Eq. 3) between observed data (laboratory analyses) and predictions generated by the 6 machine learning algorithms. The nRMSE was calculated as the ratio between RMSE and the mean of observed values (Eq. 4) and Ratio of Performance to Interquartile Distance (RPIQ) was calculated dividing the interquartile distance per the RMSE (Eq. 5). According to Chang et al. (2001), RPD values are classified as follows: RPD > 2 indicates that models provide good predictions; 1.4 ≤ RPD ≤ 2 indicates moderately good predictions; and RPD < 1.4 indicates unreliable models. Thus, good models exhibit high R² and RPD values and low RMSE. The definitions are as follows:

R 2 = R S S T S S (1)
R M S E = 1 n i = 1 n ( y i m i ) 2 (2)
R P D = S D R M S E (3)
n R M S E = R M S E M e a n (4)
R P I Q = I Q R M S E (5)

where: RSS are the sum of squares of residuals and TSS are total sum of squares. n is the number of observations, yi are the predicted values, mi are the observed values, and SD is the standard deviation of the observed values. IQ is the interquartile distance.

RESULTS

Dataset variability

The evaluated soil fertility attributes, along with their descriptive statistics, are presented in Table I. Micronutrients exhibited the highest mean coefficient of variation (CV) at 140%, whereas the groups of physical attributes and toxicity showed the lowest mean CVs at 61% and 57%, respectively. Macronutrients displayed an intermediate mean CV of 84%, while organic carbon exhibited a CV of 82%.

Table I
Descriptive statistics of the 17 soil fertility properties from the 468 surficial soil samples collected in Rio Grande do Sul state, Brazil.

The variability of the results reflects the broad range of samples with differing attributes, capturing the diversity of the sampled soils. The great range of OC may also indicate soils under different hydromorphic conditions, which tend to accumulate OC as decomposition is prevented with absence of oxygen. Exchangeable Ca also presented a very wide range (12.17 cmolc dm-3), with values from below those considered adequate for most crops, to other values surpassing the optimal contents for this nutrient.

Such range of soil fertility properties may hamper soil management and crop production in terms of achieving more homogeneous soil conditions for plant development. Conversely, heterogeneous datasets tend to be more advantageous for studies aiming to model and predict soil properties, as they encompass a wider range of soil variations. Under these circumstances, the values of pH, and available K, Zn, and Cu are comparable to those reported by Dasgupta et al. (2022) for the eastern region of India.

Predictive Models for soil fertility properties

Table II presents the performance of the six predictive models (PPR, PLS, RF, SVM, XGB, Cubist) in estimating the 17 soil fertility properties studied herein based on root mean square error (RMSE). Among the models evaluated, Cubist emerged as the one that delivered the best predictions for 59% of the properties analyzed, while RF was the most accurate for the remaining 41%.

Table II
Root mean square error resulted from the validation of the models for predicting 17 fertility properties of soils from Rio Grande do Sul state, Brazil, using 6 machine learning algorithms: PPR – projection pursuit regression, PLS – partial least squares, RF – random forest, SVM – support vector machine, XGB – extreme gradient boosting.

The Cubist algorithm was the most accurate model in predicting exchangeable/available Al, Ca, Mg, N, B, Mn, and Zn in addition to P-rem and H+Al, achieving the lowest RMSE values for these variables. Conversely, the RF model was successful in predicting available Na, Cu, Fe, K, P, S and OC. This indicates that Cubist and RF are particularly well-suited for modeling these attributes, likely due to their decision-tree-based approach, which effectively captures nonlinear patterns (Gupta et al. 2024). Also, these results suggest that both algorithms are effective in handling datasets with high variability and complex distributions, as it is a robust algorithm leveraging multiple decision trees to reduce variance.

For certain specific variables, other models also demonstrated good performance. For example, PLS exhibited relatively strong performance for available Na, with an RMSE of 4.26 g kg-1, slightly worse than the best models (RF). The same occurred with SVM for predicting OC, which delivered RMSE of 0.71 dag kg-1, the same achieved by Cubist, but both are slightly worse than that found by RF (0.65 dag kg-1). However, overall, Cubist and RF consistently outperformed the other algorithms across all the soil fertility properties analyzed.

An analysis of the predictive performance of the models (Table III) reveals a clear superiority of the Cubist and RF algorithms, with differences in their efficacy depending on the soil property. The Cubist model excelled particularly for exchangeable/available Al, Ca, H+Al, Mg, Mn, N, P_rem, and Zn, with R² values ranging from 0.52 to 0.84 and RPD values from 1.27 to 2.45. For instance, for available Mn, Cubist achieved an R² of 0.84 and an RPD of 2.45, result classified as “good” according to Chang et al. (2001).

Table III
Prediction results delivered by the best model for each soil fertility property predicted.

The Ratio of Performance to Interquartile Distance (RPIQ) was used to assess the predictive quality of models across all soil fertility attributes. Following the classification, RPIQ values were interpreted as: “excellent” (> 4.05), “good” (3.37–4.05), “approximate quantitative predictions” (2.70–3.37), and “suitable for distinguishing between high and low values” (2.02–2.70) (Saeys et al. 2005, Ludwig et al. 2017). Based on these thresholds, Cubist achieved the highest RPIQ scores, including approximate quantitative to good predictions for P-rem (3.63) and approximate quantitative predictions for, Mg, Al and H+Al with 3.07, 2.86 and 2.80 respectively. Similarly, Random Forest yielded strong results for available K (3.34), Cu (2.92) and S (2.31).

The nRMSE was used to compare prediction accuracy across soil attributes with different scales. According to the classification proposed by Mozaffari et al. (2024a), values below 10% indicate excellent, 10–20% good, 20–30% fair, and above 30% poor predictions. In this study, nRMSE values ranged from 0.08% to 1.62%, with the best performances observed for pH (0.08%), available K (0.17%), and P-rem (0.19%), indicating excellent predictions. These results are consistent with R² and RPIQ trends, reinforcing the models varying capacity to generalize across different chemical elements. The scatter plots between observed and predicted values for all the best models per soil fertility attribute predicted are presented in Figure 2.

Figure 2
Observed and predicted values delivered by the best models (Table 3) per soil fertility attribute predicted.

The analysis of macronutrients (Ca, K, Mg, N, P, P-Rem, and S) shows considerable variation in both R² and RPD values (Figure 3). Available K and P-Rem stand out with the highest RPD values, exceeding 2.0, indicating high precision and reliability in the model’s estimates for these variables. However, other macronutrients, such as N and P, exhibited moderate R² values, around 0.5, suggesting that while predictions are reasonably accurate, there is room for improvement. The RPD values for macronutrients like exchangeable Ca and Mg are also notable (1.51 and 1.97, respectively), demonstrating moderate robustness in the predictions of these essential soil attributes.

Figure 3
Performance of predictive models in estimating different groups of soil chemical attributes, based on R² and RPD values.

Among the micronutrients (B, Cu, Fe, Mn, and Zn), even greater variability is observed. The models showed high accuracy for Cu and Mn, with RPDs exceeding 2.0, indicating excellent predictive performance. The R² values for these elements were also high, particularly for available Cu, reflecting the model’s capacity to deliver accurate and reliable estimates. On the other hand, Fe, despite performing well in terms of RPD, exhibited a moderate R² value (0.46), which may indicate that while the predictions are reliable, the variability explained by the model could still be improved.

The acidity-related attributes (Al, H+Al, and pH) presented interesting variations. pH, crucial for various chemical reactions in the soil, demonstrated relatively good performance in terms of R², indicating that the model effectively explains the observed variability. However, its RPD was lower, suggesting that prediction precision could be further improved. For exchangeable Al and H+Al, both R² and RPD values were moderate, indicating that the model is reasonably effective in predicting soil toxicity.

Both OC and available Na were evaluated separately and showed distinct performances. Available Na had the lowest R² (0.16) and RPD (1.02) values, indicating that the model is not reliable. Conversely, OC displayed moderate performance, with reasonable R² (0.49) and RPD (1.39) values, suggesting that predictions for organic carbon were more accurate. OC and available Na showed moderate to low RMSE values, indicating acceptable performance of the predictive models for these variables. The analysis of RMSE values thus reveals significant variations in the accuracy of models across different fertility properties, identifying areas where the models perform well and those requiring improvements.

Figure 4 illustrates the performance of predictive models in estimating groups of soil fertility properties based on RMSE values. For macronutrients (Ca, K, Mg, N, P, P-Rem, and S), available K prediction exhibited the highest RMSE, indicating greater difficulty in precisely estimating this nutrient. Conversely, exchangeable Mg and total N showed relatively low RMSE values, suggesting that the models are more effective in predicting these nutrients.

Figure 4
Performance of predictive models in estimating groups of soil chemical attributes based on RMSE values.

For micronutrients (B, Cu, Fe, Mn, and Zn), available Fe stands out with a significantly high RMSE, highlighting challenges in predictive modeling for this specific element using the pXRF-based dataset. In contrast, available B and Cu exhibited lower RMSE values, suggesting greater model precision in estimating these variables. Regarding acidity-related properties, exchangeable Al and H+Al showed similar and relatively low RMSE values, indicating good model performance in predicting these attributes. The pH also presented moderate RMSE values, indicating that the models can adequately estimate this critical soil variable.

Importance of pXRF variables for soil fertility predictions

The importance of pXRF predictor covariates (Figure 5) in estimating soil fertility properties reveals interesting patterns in soil attribute prediction, highlighting the relevance of specific elements measured by pXRF. Notably, silicon (Si) appeared 13 times as a crucial covariate, followed by zinc (Zn) with 12 occurrences and iron (Fe) with 11 occurrences among the five most important variables for the models.

Figure 5
Importance of pXRF covariates in estimating soil chemical attributes.

The analysis of variable importance revealed that specific pXRF-detected elements contributed differently to the prediction of each soil fertility property. All models were built using the same set of 13 elemental variables obtained by pXRF (Al, Ca, Fe, K, Mg, P, Rb, Si, Sr, Ti, V, Zn, and Zr). From them, certain elements appeared more frequently among the top five most important predictors across the target soil attributes. As shown in Figure 5, Rb, Zn, Fe, Si, and P detected by pXRF were among the most recurrent top predictor variables, highlighting their broad predictive relevance. The top-ranked predictor (1st position) varied according to the target attribute. For example, Ca was the most important for predicting exchangeable Ca and available B, while P stood out as the primary covariate for predicant available P and Na, total N, and OC. This ranking offers a clear overview of the most influential pXRF elements across different soil fertility attributes.

Exploring the importance of covariates and model performance, elements like Si, Zn, Fe, and Rb, which frequently appear as important covariates, align with the variables for which models like Cubist and RF demonstrated strong performance in terms of RPD. For instance, in the case of the prediction of available Cu, better predicted by Cubist, the most important predictor variables included the top four most frequently appearing covariates, while the fifth most important one was more specific for this soil property prediction (Fe). A similar pattern was observed for available K prediction based on pXRF-detected K.

DISCUSSION

Dataset Variability

All groups and individual soil fertility properties displayed a wide range of variation (Table 1), which may be associated with differences in soil parent materials, history of land use or other soil-related features, such as topography and soil weathering degree (Hengl et al. 2017, Resende et al. 2014). Among the macronutrients, N and P show positive skewness, indicating distributions with values higher than the mean and suggesting a more significant presence of these nutrients in the soil. Probably, this is a consequence of continuous fertilizer applications over the years of eucalyptus cultivation (Xu et al. 2020, Yesilonis et al. 2016). For micronutrients, available Fe and Mn also exhibit positive skewness, while available Cu shows negative skewness, pointing to lower values. These distribution patterns can impact plant health and growth (Marschner 2012).

Regarding acidity-related properties, soil pH, which influences the solubility of many nutrients (Mascarenhas et al. 2013), shows positive skewness, suggesting a tendency toward higher values. Conversely, exchangeable Al and H+Al exhibit different distribution patterns, reflecting potential variations in soil acidity and cation exchange capacity across study areas.

We observed that the mean values of available/exchangeable Ca, K, Mg, P, as well as pH and H+Al differ from the dataset of Fontenelli et al. (2021), which included 423 samples from São Paulo State, southeastern Brazil. In the current dataset, the mean values of available/exchangeable Ca, K, and Mg are higher than those reported by Fontenelli et al. (2021), probably as a consequence of the varying soil parent materials comparing those from São Paulo state and the ones evaluated in RS soils, in addition to the management of fertilizers adopted. Additionally, the mean pH is lower in the current dataset, indicating more acidic soils. The mean H+Al value is higher in this study, suggesting greater potential acidity. The mean P content is slightly higher in the current dataset, indicating comparable soil fertility to the study by Fontenelli et al. (2021).

In Andrade et al. (2021), which utilized a dataset of 1,514 samples from seven states in southern, southeastern, and northeastern Brazil, the mean values of available B (0.15 mg kg-1) and Mn (51.57 0.15 mg kg-1) were higher than those found in this study (0.05 mg kg-1 for available B and 22.24 0.15 mg kg-1 for available Mn). The mean values of available Cu and Zn also differ, with Andrade et al. (2021) reporting a mean available Cu content of 1.19 0.15 mg kg-1 compared to 0.75 0.15 mg kg-1 in this study. For available Zn, the mean is 3.33 0.15 mg kg-1 in Andrade et al. (2021) compared to 1.98 0.15 mg kg-1 in this study. Conversely, the mean available Fe content is higher in this study at 139.19 0.15 mg kg-1, compared to 91.40 0.15 mg kg-1 in Andrade et al. (2021). More importantly, soil fertility properties are affected by several soil-, climate-, management-related factors (Lopes & Guilherme 2016), which explains the variation of ranges of those properties.

Predictions of soil fertility properties

The results presented in Table 2 indicate that the Cubist and RF models are highly recommended for predicting soil fertility properties due to their high precision and robustness, as evidenced by the lower RMSE compared to the ones delivered by the other 5 algorithms tested. These results are consistent with those reported by Chatterjee et al. (2021) and O’Rourke et al. (2016), also achieving better predictions with those algorithms. The adoption of these models can significantly contribute to more efficient and sustainable agricultural practices, enabling detailed and accurate analysis of soil conditions, in addition to reducing costs for soil characterization (Benedet et al. 2021a).

A comparative analysis of model performance for soil fertility predictions (Table 3) shows that the Cubist and RF algorithms consistently outperformed the other models tested. The Cubist model stood out particularly for attributes such as available/exchangeable Al, Ca, Mg, Mn, N, and Zn, as well as H+Al and P_Rem, with R² values ranging from 0.52 to 0.84 and RPD values from 1.27 to 2.45. For instance, for available Mn, Cubist achieved an R² of 0.84 and an RPD of 2.45, classifying it as “good” according to Chang et al. (2001). Additionally, Cubist was effective in predicting available B and pH, with R² values of 0.47 and 0.56 and RPD values of 1.37 and 1.51, respectively. While these R² and RPD values are moderate, they still represent better performance compared to other models for these variables, suggesting that Cubist may still be preferable for predictions of available B and pH due to its overall consistency. These high R² values indicate that Cubist effectively explains the observed data variability, while the high RPD values suggest it is reliable and useful for predicting these variables.

Random Forest (RF) also demonstrated significant performance, particularly for OC, and available Cu, Fe, K, Na, P, and S. With R² values of 0.49 for OC, 0.81 for available Cu, 0.46 for available Fe, 0.82 for available K, 0.16 for available Na, 0.54 for available P, and 0.60 for available S, RF exhibits strong explanatory capacity, particularly notable for available Cu and K. The RPD values of 2.24 for available Cu and 2.32 for available K classify these models as “good,” indicating that RF not only explains a high proportion of data variability but also provides reliable predictions for these variables. The results for available Cu and Mn are more accurate than those achieved by Pelegrino et al. (2019) (R² = 0.63) also using RF and pXRF data, indicating the efficiency of prediction models may be local/regional-specific (Faria et al. 2020) and/or attribute-specific (Tavares et al. 2020a).

The RPIQ-based classification (Table 3) provided a deeper understanding of the predictive potential of each algorithm beyond standard metrics such as R² and RMSE. The Cubist model showed the most consistent high-level performance, with RPIQ values indicating “approximate quantitative” to “good” predictive capability for P-rem (3.63), and exchangeable Mg (3.07) and Al (2.86), as well as for H+Al (2.80). These results confirm the model’s robustness in capturing complex patterns in fertility-related attributes, especially those related to soil acidity and P dynamics. Random Forest also showed strong predictive capacity, particularly for available K (3.34), Cu (2.92), and S (2.31), falling into the “approximate quantitative” category. These findings align with previous studies (Benedet et al. 2021b, Karami et al. 2024a) that reported superior performance of ensemble-based models in soil prediction tasks due to their ability to handle non-linearity, variable interactions, and data heterogeneity.

The use of nRMSE complemented other evaluation metrics by allowing standardized comparison across soil fertility variables of distinct magnitudes. The Cubist and Random Forest models achieved excellent predictive accuracy (nRMSE < 10%) for available Mn and Cu, and exchangeable Al, confirming their reliability for operational soil assessment. Similar findings were reported by Mozaffari et al. (2024a) and Karami et al. (2024b), who emphasized that high nRMSE values are often observed in elements present in trace amounts or with narrow concentration ranges in soil datasets. The consistent performance of Cubist and RF models highlights their suitability for operational use in pXRF-driven fertility assessment frameworks.

The superior performance of the Cubist model in predicting several soil fertility properties can be attributed to its unique structure, which combines regression trees with linear models at the terminal nodes (Gozukara et al. 2022b). This hybrid approach allows Cubist to partition the data into homogeneous subsets while also capturing linear relationships within each partition, making it particularly effective for modeling complex and heterogeneous soil datasets. Unlike purely tree-based models such as Random Forest, which aggregate multiple decision trees without internal linear modeling, Cubist leverages both rule-based segmentation and multivariate regression. This makes it especially suitable for prediction of soil attributes from diverse physiographic regions, where different processes may govern the behavior of the same variable under varying conditions (Teixeira et al. 2022). Additionally, the ability of Cubist to generate interpretable rules enhances its utility in understanding the covariate structure, while its internal model averaging helps reduce overfitting, often observed in high-dimensional environmental data. These characteristics likely contributed to the consistently high R² and low RMSE values observed across multiple soil fertility attributes in this study.

Good performance of the Cubist and RF algorithms for predicting soil fertility attributes reinforces the potential of combining machine learning with proximal sensing to support rapid soil diagnostics. These findings align with previous studies demonstrating that ensemble and rule-based models tend to outperform traditional linear regressions, particularly in heterogeneous soil datasets (Zhang & Hartemink 2020, Padarian et al. 2017). For example, Gozukara et al. (2022b) observed similar success using Cubist to predict soil texture and chemical attributes from Vis-NIR data under variable field conditions. Our study expands on this by validating the use of pXRF in a subtropical environment with highly weathered soils, a condition less explored in global literature. Furthermore, the ability to predict 12 out of 17 fertility attributes with acceptable accuracy has significant implications for sustainable land management. It offers a pathway for reducing dependency on chemical analyses, lowering costs, and supporting faster decision-making in precision agriculture (Feng et al. 2021).

The comparatively lower performance of PPR, PLS, SVM, and XGB can be attributed to both algorithm-specific limitations and the inherent complexity of soil datasets. PLS and PPR are linear or semi-linear models that tend to underperform when capturing non-linear relationships between elemental data and fertility attributes (Tavares et al. 2020a). SVM, although capable of modeling non-linearities through kernel functions, is highly sensitive to parameter tuning and data scaling, which can limit its effectiveness in heterogeneous datasets with non-normal distributions (Keskin et al. 2019). XGB, while robust in many applications, may overfit when dealing with small to medium-sized datasets that contain noise or strong multicollinearity (Zhao et al. 2022).

PXRF offers a comprehensive view of total elemental composition in soils, while conventional analyses for soil fertility assessment focus on nutrients promptly available to plants (Fischer et al. 2020). This point hampers the direct relationships between pXRF results and soil fertility properties (Rawal et al. 2019). However, these relationships tend to work better for some elements that may be less commonly found in the crystalline structure of minerals or present in soil organic matter, which was the case of Ca in the Brazilian cerrado soils (Teixeira et al. 2018). Conversely, K and Al are commonly found in the crystalline structure of minerals widely spread in Brazilian soils, such as muscovite (KAl2(Si3Al)O10(OH,F)2) and kaolinite (Al2Si2O5(OH)4) (Kämpf et al. 2012), which prevents the direct relationships between pXRF and soil fertility data. In such cases, the association of the other elements detected by pXRF help explain soil fertility and other properties mostly via machine learning algorithms (Xu et al. 2019, Gozukara et al. 2022a, Tavares et al. 2025, Teixeira et al. 2022).

As shown in Figure 5, the prominence of Si and Zn detected by pXRF suggests they are robust indicators for a wide range of soil fertility properties. Additionally, the variety of pXRF elements that emerge as important across different contexts indicates the complexity of interactions between soil fertility properties and covariates (Andrade et al. 2020a, Antonangelo & Zhang 2024). Elements such as Rb and Al, which appear multiple times but less frequently than Si and Zn, still play critical roles in specific contexts. For example, Rb is the most important covariate for exchangeable Al and available K predictions, indicating that despite its lower overall frequency, its impact should not be underestimated in specific analyses.

The recurrence of certain pXRF-derived elements as important predictors reflects underlying geochemical and pedogenetic processes. The high relevance of Rb may be associated with its role as a tracer of mineral weathering and parent material, contributing to the prediction of exchangeable Al and available Cu, K, and Mn (Tóth et al. 2019, Gjengedal et al. 2015). Similarly, Fe and Si detected by pXRF were closely related to pH and H+Al, likely due to their influence on mineral buffering capacity and soil acidity regulation in weathered tropical soils (Weindorf et al. 2014, Simonsson et al. 2016). Zn detected by pXRF was also consistently important, especially for micronutrient predictions, possibly due to its co-occurrence with oxides and involvement in nutrient sorption mechanisms (Duffner et al. 2014). These associations support the interpretation that total elemental data from pXRF can serve as effective proxies for both direct nutrient contents and indirect indicators of soil fertility dynamics, enhancing model accuracy and interpretability.

Interestingly, the element directly measured by pXRF is not always the primary predictor of the content of this same element in the available/exchangeable form. For instance, in available Cu estimation, elements such as Rb and Zn measured by pXRF are more important than Cu itself, a relationship observed in other similar studies (Andrade et al. 2021). This discrepancy suggests that correlations between total element contents measured by pXRF and available contents can be complex, reflecting soil chemical and physical interactions that influence nutrient availability. The frequent importance of Zn and Fe detected by pXRF as predictors for various available/exchangeable contents suggests that these total elements captured by pXRF may better reflect the complex soil dynamics than their available counterparts.

Thus, using pXRF to measure total elements with the aid of machine learning methods to detect a relationship between them and soil fertility properties proves to be a valuable approach. Understanding these relationships is essential to optimize modeling and prediction of soil chemical attributes, enabling the selection of the most influential covariates to improve accuracy and reliability. Moreover, this approach also allows for new discoveries about the relationships between total contents and soil fertility properties little investigated so far (Silva et al. 2020), such as the importance of Rb for prediction of exchangeable Al and available K in the soils here evaluated.

CONCLUSIONS

The evaluated soils of Rio Grande do Sul state, Brazil, exhibited significant variability in their chemical properties, influenced by diverse influenced by diverse geological, pedological and climatic factors and climatic factors, as they are submitted to the same management practices regarding fertilization. This study demonstrated that portable X-ray fluorescence (pXRF), when combined with machine learning algorithms, is a viable and efficient tool for predicting multiple soil fertility attributes under the diverse pedoclimatic conditions. Among the tested algorithms, Cubist and Random Forest showed the highest predictive performance, with Cubist excelling for exchangeable/available Al, Ca, Mg, Mn, and Zn, as well as total N, potential acidity (H+Al), and P-rem (R² up to 0.84), while Random Forest performed better for available Cu, K, P, and S.

Out of the 17 soil fertility attributes, 12 were predicted with acceptable to good accuracy, reinforcing the potential of pXRF for operational use in soil fertility diagnosis. Elements such as Zn, Fe, Si, and Rb detected by pXRF emerged as key predictors. However, limitations were observed for predicting organic carbon and available contents of B, Fe, Na, and Zn, indicating the need for further investigations. The results delivered by pXRF compose a rich database that can be used along with machine learning algorithms to accurately estimate available soil nutrient contents and other soil fertility properties. Although this study was conducted under specific pedoclimatic conditions of Rio Grande do Sul, the methodological framework used here involving pXRF data integrated with supervised machine learning algorithms is broadly applicable to other regions after model calibration and validation. However, successful transferability of the models developed herein to other areas requires validation of the predictions at those areas followed by fine adjustments of those models.

Future research should aim to expand pXRF-based modeling to include a wider range of soil classes and management systems, especially in tropical and subtropical regions with even more contrasting soils. The integration of pXRF data into national-scale digital soil mapping frameworks also holds promise for enhancing soil fertility assessment at more detailed spatial scales, contributing to precision management and sustainable land-use planning.

Acknowledgements

The authors would like to thank Conselho Nacional de Desenvolvimento Científico e Tecnológico (CNPq, 404983/2023-5; 307532/2022-4), Coordenação de Aperfeiçoamento Pessoal de Nível Superior (CAPES), Fundação de Amparo à Pesquisa do Estado de Minas Gerais (FAPEMIG, APQ-02907-18), and Serviço Geológico Brasileiro (SGB) TED UFLA/CPRM number 241/2021 for the scholarships and other fundings provided to the development of this study. The authors also acknowledge the National Institute of Science and Technology on Soil and Food Security (CNPq grant 406577/2022-6) for its financial support for this work.

References

  • ALVARES CA, STAPE JL, SENTELHAS PC & DE MORAES GONÇALVES JL. 2013a. Modeling monthly mean air temperature for Brazil. Theor Appl Climatol 113: 407-427.
  • ALVARES CA, STAPE JL, SENTELHAS PC, DE MORAES GONÇALVES JL & SPAROVEK G. 2013b. Köppen’s climate classification map for Brazil. Meteorol Zeitschrift 22: 711-728.
  • ALVAREZ VVH, NOVAIS R, DIAS L & OLIVEIRA JA. 2000. Determinação e uso do fósforo remanescente. B Inf SBCS 25: 27-32.
  • ANDRADE R, FARIA WM, SILVA SHG, CHAKRABORTY S, WEINDORF DC, MESQUITA LF, GUILHERME LRG & CURI N. 2020a. Prediction of soil fertility via portable X-ray fluorescence (pXRF) spectrometry and soil texture in the Brazilian Coastal Plains. Geoderma 357: 113960.
  • ANDRADE R ET AL. 2023. Proximal sensing provides clean, fast, and accurate quality control of organic and mineral fertilizers. Environ Res 236. https://doi.org/10.1016/j.envres.2023.116753
    » https://doi.org/10.1016/j.envres.2023.116753
  • ANDRADE R, SILVA SHG, FARIA WM, POGGERE GC, BARBOSA JZ, GUILHERME LRG & CURI N. 2020b. Proximal sensing applied to soil texture prediction and mapping in Brazil. Geoderma Reg 23. https://doi.org/10.1016/j.geodrs.2020.e00321
    » https://doi.org/10.1016/j.geodrs.2020.e00321
  • ANDRADE R, SILVA SHG, WEINDORF DC, CHAKRABORTY S, FARIA WM, GUILHERME LRG & CURI N. 2021. Micronutrients prediction via pXRF spectrometry in Brazil: Influence of weathering degree. Geoderma Reg 27: e00431.
  • ANTONANGELO J & ZHANG H. 2024. Assessment of portable X-ray fluorescence (pXRF) for plant-available nutrient prediction in biochar-amended soils. Sci Rep 14: 20377.
  • BAGATINI T, GIASSON E & TESKE R. 2016. Expansão de mapas pedológicos para áreas fisiograficamente semelhantes por meio de mapeamento digital de solos. Pesqui Agropecu Bras 51: 1317-1325.
  • BELLON-MAUREL V, FERNANDEZ-AHUMADA E, PALAGOS B, ROGER J-M & MCBRATNEY A. 2010. Critical review of chemometric indicators commonly used for assessing the quality of the prediction of soil attributes by NIR spectroscopy. TrAC Trends Anal Chem 29: 1073-1081.
  • BENEDET L ET AL. 2021a. Rapid soil fertility prediction using X-ray fluorescence data and machine learning algorithms. CATENA 197: 105003.
  • BENEDET L, NILSSON MS, SILVA SHG, PELEGRINO MHP, MANCINI M, DE MENEZES MD, GUILHERME LRG & CURI N. 2021b. X-ray fluorescence spectrometry applied to digital mapping of soil fertility attributes in tropical region with elevated spatial variability. An Acad Bras Cienc 93: e20200646. https://doi.org/10.1590/0001-3765202120200646.
    » https://doi.org/10.1590/0001-3765202120200646
  • BENNETT KP & CAMPBELL C. 2000. Support vector machines. ACM SIGKDD Explor Newsl 2: 1-13.
  • BOEHMKE B & GREENWELL B. 2019. Random Forests. In: Hands-On Machine Learning with R, Chapman and Hall/CRC, p. 203-219.
  • BONFATTI BR, HARTEMINK AE, GIASSON E, TORNQUIST CG & ADHIKARI K. 2016. Digital mapping of soil carbon in a viticultural region of Southern Brazil. Geoderma 261. https://doi.org/10.1016/j.geoderma.2015.07.016
    » https://doi.org/10.1016/j.geoderma.2015.07.016
  • BREIMAN L. 2001. Random Forests. Mach Learn 45: 5-32.
  • BREIMAN L, CUTLER A, LIAW A & WIENER M. 2002. randomForest: Breiman and Cutlers Random Forests for Classification and Regression. CRAN Contrib Packag 103: 239-248.
  • BRUNGARD C, NAUMAN T, DUNIWAY M, VEBLEN K, NEHRING K, WHITE D, SALLEY S & ANCHANG J. 2021. Regional ensemble modeling reduces uncertainty for digital soil mapping. Geoderma 397: 114998.
  • CHANG S-T, WU J-H, WANG S-Y, KANG P-L, YANG N-S & SHYUR L-F. 2001. Antioxidant Activity of Extracts from Acacia confusa Bark and Heartwood. J Agric Food Chem 49: 3420-3424.
  • CHATTERJEE S, HARTEMINK AE, TRIANTAFILIS J, DESAI AR, SOLDAT D, ZHU J, TOWNSEND PA, ZHANG Y & HUANG J. 2021. Characterization of field-scale soil variation using a stepwise multi-sensor fusion approach and a cost-benefit analysis. Catena 201: 105190.
  • CHEN T & GUESTRIN C. 2016. XGBoost. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, New York, NY, USA: ACM, p.785-794.
  • CHEN T ET AL. 2025. xgboost: Extreme Gradient Boosting. CRAN Contrib Packag.
  • DASGUPTA S, CHAKRABORTY S, WEINDORF DC, LI B, SILVA SHG & BHATTACHARYYA K. 2022. Influence of auxiliary soil variables to improve PXRF-based soil fertility evaluation in India. Geoderma Reg 30: e00557.
  • DUFFNER A, WENG L, HOFFLAND E & VAN DER ZEE SEATM. 2014. Multi-surface Modeling To Predict Free Zinc Ion Concentrations in Low-Zinc Soils. Environ Sci Technol 48: 5700-5708.
  • FAO, IFAD, UNICEF, WFP & WHO. 2023. The State of Food Security and Nutrition in the World 2023, 1-316 p.
  • FARIA ÁJG, SILVA SHG, MELO LCA, ANDRADE R, MANCINI M, MESQUITA LF, DOS SANTOS TEIXEIRA AF, GUILHERME LRG & CURI N. 2020. Soils of the Brazilian Coastal Plains biome: prediction of chemical attributes via portable X-ray fluorescence (pXRF) spectrometry and robust prediction models. Soil Res 58: 683-695.
  • FENG X, ZHANG H & YU P. 2021. X-ray fluorescence application in food, feed, and agricultural science: a critical review. Crit Rev Food Sci Nutr 61: 2340-2350.
  • FISCHER S, HILGER T, PIEPHO H-P, JORDAN I, KARUNGI J, TOWETT E, SHEPHERD K & CADISCH G. 2020. Soil and farm management effects on yield and nutrient concentrations of food crops in East Africa. Sci Total Environ 716: 137078.
  • FONTENELLI JV, ADAMCHUK VI, FERREIRA MMC, AMARAL LR, GUIMARÃES CCB, DEMATTÊ JAM & MAGALHÃES PSG. 2021. Evaluating the synergy of three soil spectrometers for improving the prediction and mapping of soil properties in a high anthropic management area: A case of study from Southeast Brazil. Geoderma 402: 115347.
  • FRIEDMAN JH. 2001. Greedy Function Approximation: A Gradient Boosting Machine. Ann Stat 29: 1189-1232.
  • FRIEDMAN JH & STUETZLE W. 1981. Projection Pursuit Regression. J Am Stat Assoc 76: 817-823.
  • GJENGEDAL E, MARTINSEN T & STEINNES E. 2015. Background levels of some major, trace, and rare earth elements in indigenous plant species growing in Norway and the influence of soil acidification, soil parent material, and seasonal variation on these levels. Environ Monit Assess 187: 386.
  • GOZUKARA G, ALTUNBAS S, DENGIZ O & ADAK A. 2022a. Assessing the effect of soil to water ratios and sampling strategies on the prediction of EC and pH using pXRF and Vis-NIR spectra. Comput Electron Agric 203: 107459.
  • GOZUKARA G, ZHANG Y & HARTEMINK AE. 2022b. Using pXRF and vis-NIR spectra for predicting properties of soils developed in loess. Pedosphere 32: 602-615.
  • GUPTA S, HASLER JK & ALEWELL C. 2024. Mining soil data of Switzerland: New maps for soil texture, soil organic carbon, nitrogen, and phosphorus. Geoderma Reg 36: e00747.
  • HENGL T ET AL. 2017. Soil nutrient maps of Sub-Saharan Africa: assessment of soil nutrient content at 250 m spatial resolution using machine learning. Nutr Cycl Agroecosystems 109: 77-102.
  • KÄMPF N, MARQUES J & CURI N. 2012. Mineralogia de solos brasileiros. In: Ker JC et al. (Eds), Pedologia Fundamentos, Viçosa: Sociedade Brasileira de Ciência do Solo, p. 81-146.
  • KARAMI A, MOOSAVI AA, POURGHASEMI HR, RONAGHI A, GHASEMI-FASAEI R & LADO M. 2024a. Application of proximal sensing approach to predict cation exchange capacity of calcareous soils using linear and nonlinear data mining algorithms. J Soils Sediments 24: 2248-2267.
  • KARAMI A, MOOSAVI AA, POURGHASEMI HR, RONAGHI A, GHASEMI-FASAEI R, VIDAL E & LADO M. 2024b. Proximal sensing approach for characterization of calcareous soils using multiblock data analysis. Geoderma Reg 36: e00752.
  • KESKIN H, GRUNWALD S & HARRIS WG. 2019. Digital mapping of soil carbon fractions with machine learning. Geoderma 339. https://doi.org/10.1016/j.geoderma.2018.12.037
    » https://doi.org/10.1016/j.geoderma.2018.12.037
  • KUHN M. 2008. Building Predictive Models in R Using the caret Package. J Stat Softw 28: 1-26.
  • KUHN M & QUINLAN R. 2011. Cubist: Rule- And Instance-Based Regression Modeling. CRAN Contrib Packag.
  • LILAND KH, MEVIK B-H & WEHRENS R. 2024. pls: Partial Least Squares and Principal Component Regression. CRAN Contrib Packag.
  • LIMA RF, APARECIDO LEO, TORSONI GB & ROLIM GS. 2023. Climate Change Assessment in Brazil: Utilizing the Köppen-Geiger (1936) Climate Classification. Rev Bras Meteorol 38: 1-33.
  • LIU F, WU H, ZHAO Y, LI D, YANG J-L, SONG X, SHI Z, ZHU A-X & ZHANG G-L. 2022. Mapping high resolution National Soil Information Grids of China. Sci Bull 67: 328-340.
  • LOPES AS & GUIMARÃES GUILHERME LR. 2016. Chapter One - A Career Perspective on Soil Management in the Cerrado Region of Brazil. In: Sparks DLBT-A (Ed), Advances in Agronomy, Academic Press, p. 1-72.
  • LUDWIG B, VORMSTEIN S, NIEBUHR J, HEINZE S, MARSCHNER B & VOHLAND M. 2017. Estimation accuracies of near infrared spectroscopy for general soil properties and enzyme activities for two forest sites along three transects. Geoderma 288: 37-46.
  • MANCINI M ET AL. 2024. Multinational prediction of soil organic carbon and texture via proximal sensors. Soil Sci Soc Am J 88: 8-26.
  • MARSCHNER H. 2012. Marschner’s mineral nutrition of higher plants, Academic press.
  • MASCARENHAS HAA, ESTEVES JAF, WUTKE EB, RECO PC & LEÃO PCL. 2013. Disability and toxicity of nutrients in visual soybeans. Nucleus 10: 281-306.
  • MCLEAN EO, HEDDLESON MR, BARTLETT RJ & HOLOWAYCHUK N. 1958. Aluminum in Soils: I. Extraction Methods and Magnitudes in Clays and Ohio Soils. Soil Sci Soc Am J 22: 382-387.
  • MEHLICH A. 1953. Determination of P, Ca, Mg, K, Na and NH4 by North Carolina soil testing laboratories. Raleigh, Univ North Carolina.
  • MEYER D, DIMITRIADOU E, HORNIK K, WEINGESSEL A & LEISCH F. 2024. e1071: Misc Functions of the Department of Statistics, Probability Theory Group (Formerly: E1071), TU Wien. CRAN Contrib Packag.
  • MOURA-BUENO JM, DALMOLIN RSD, HORST-HEINEN TZ, CANCIAN LC, SCHENATO RB, DOTTO AC & FLORES CA. 2019. Prediction of soil classes in a complex landscape in Southern Brazil. Pesqui Agropecuária Bras 54.
  • MOURA-BUENO JM, SIMÃO R, DALMOLIN D, HORST-HEINEN Z, TEN CATEN A, VASQUES GM, DOTTO AC & GRUNWALD S. 2020. When does stratification of a subtropical soil spectral library improve predictions of soil organic carbon content?
  • MOZAFFARI H, MOOSAVI AA, BAGHERNEJAD M & CORNELIS W. 2024a. Revisiting soil texture analysis: Introducing a rapid single-reading hydrometer approach. Measurement 228: 114330.
  • MOZAFFARI H, MOOSAVI AA & OSTOVARI Y. 2024b. Feasibility of proximal sensing for predicting soil loss tolerance. Catena 247: 108503.
  • O’ROURKE SM, STOCKMANN U, HOLDEN NM, MCBRATNEY AB, MINASNY B. 2016. An assessment of model averaging to improve predictive power of portable vis-NIR and XRF for the determination of agronomic soil properties. Geoderma 279: 31-44.
  • PADARIAN J, MINASNY B & MCBRATNEY AB. 2017. Chile and the Chilean soil grid: A contribution to GlobalSoilMap. Geoderma Reg 9: 17-28.
  • PELEGRINO MHP, SILVA SHG, DE FARIA ÁJG, MANCINI M, TEIXEIRA AFS, CHAKRABORTY S, WEINDORF DC, GUILHERME LRG & CURI N. 2022. Prediction of soil nutrient content via pXRF spectrometry and its spatial variation in a highly variable tropical area. Precis Agric 23: 18-34.
  • PELEGRINO MHP, SILVA SHG, MENEZES MD, SILVA E, OWENS PR & CURI N. 2016. Mapping soils in two watersheds using legacy data and extrapolation for similar surrounding areas. Ciência e Agrotecnologia 40: 534-546.
  • PELEGRINO MHP, WEINDORF DC, SILVA SHG, DE MENEZES MD, POGGERE GC, GUILHERME LRG & CURI N. 2019. Synthesis of proximal sensing, terrain analysis, and parent material information for available micronutrient prediction in tropical soils. Precis Agric 20: 746-766.
  • RAWAL A, CHAKRABORTY S, LI B, LEWIS K, GODOY M, PAULETTE L & WEINDORF DC. 2019. Determination of base saturation percentage in agricultural soils via portable X-ray fluorescence spectrometer. Geoderma 338: 375-382.
  • RESENDE M, CURI N, RESENDE SB, CORRÊA GF & KER JC. 2014. Pedologia Base Para Distinção de Ambientes, 6ª ed., Lavras: UFLA, 378 p.
  • RIBEIRO JV, DOS SANTOS FR, DE OLIVEIRA JF, BARBOSA GMC & MELQUIADES FL. 2024. Optimization of pXRF instrumentation conditions and multivariate modeling in soil fertility attributes determination. Spectrochim Acta Part B At Spectrosc 211: 106835.
  • SAEYS W, MOUAZEN AM & RAMON H. 2005. Potential for onsite and online analysis of pig manure using visible and near infrared reflectance spectroscopy. Biosyst Eng 91: 393-402.
  • SILVA SHG ET AL. 2020. Soil texture prediction in tropical soils: A portable X-ray fluorescence spectrometry approach. Geoderma 362: 114136.
  • SILVA SHG ET AL. 2021. pXRF in tropical soils: Methodology, applications, achievements and challenges. Adv Agron 167: 1-62.
  • SIMONSSON M, COURT M, BERGHOLM J, LEMARCHAND D & HILLIER S. 2016. Mineralogy and biogeochemistry of potassium in the Skogaby experimental forest, southwest Sweden: pools, fluxes and K/Rb ratios in soil and biomass. Biogeochemistry 131: 77-102.
  • SMOLA AJ & SCHÖLKOPF B. 2004. A tutorial on support vector regression. Stat Comput 14: 199-222.
  • STRECK EV, KAMPF N, DALMOLIN S, KLAMT E, NASCIMENTO PC, SCHNEIDER P, GIASSON E & PINTO LFS. 2018. Solos do Rio Grande do Sul, 3ª ed., Porto Alegre: Emater/RS-Ascar, 252 p.
  • TANURE TMP, DOMINGUES EP & MAGALHÃES AS. 2024. Regional impacts of climate change on agricultural productivity: evidence on large-scale and family farming in Brazil. Rev Econ Sociol Rural 62.
  • TAVARES TR, MINASNY B, MCBRATNEY A, MOLIN JP, MARQUES GT, RAGAGNIN MM, DOS SANTOS FR, DE CARVALHO HWP & LAVRES J. 2025. Do XRF local models have temporal stability for predicting plant-available nutrients in different years? A long-term study showing the effect of soil fertility management in a tropical field. Soil Tillage Res 245: 106307.
  • TAVARES TR, MOLIN JP, JAVADI SH, CARVALHO HWP & MOUAZEN AM. 2020a. Combined Use of Vis-NIR and XRF Sensors for Tropical Soil Fertility Analysis: Assessing Different Data Fusion Approaches. Sensors 21: 148.
  • TAVARES TR, MOUAZEN AM, ALVES EEN, DOS SANTOS FR, MELQUIADES FL, PEREIRA DE CARVALHO HW & MOLIN JP. 2020b. Assessing soil key fertility attributes using a portable X-ray fluorescence: A simple method to overcome matrix effect. Agronomy 10: 787.
  • TEIXEIRA AFS, ANDRADE R, MANCINI M, SILVA SHG, WEINDORF DC, CHAKRABORTY S, GUILHERME LRG & CURI N. 2022. Proximal sensor data fusion for tropical soil property prediction: Soil fertility properties. J South Am Earth Sci 116: 103873.
  • TEIXEIRA AFS, WEINDORF DC, SILVA SHG, GUILHERME LRG & CURI N. 2018. Portable x-ray fluorescence (pXRF) spectrometry applied to the prediction of chemical attributes in inceptisols under different land use. Cienc e Agrotecnologia 42: 501-512.
  • TEIXEIRA PC, DONAGEMMA GK, FONTANA A & TEIXEIRA WG. 2017. Manual de métodos de análise de solo, 3ª ed., Brasília, DF: Embrapa Solos, 574 p.
  • TORNQUIST CG, GAMBOA CH, ANDRIOLLO DD, REICHERT JM & DOS SANTOS FJ. 2024. Soil Carbon Stocks in the Brazilian Pampa: An Update. In: Overbeck GE et al. (Eds), South Brazilian Grasslands, Cham: Springer International Publishing, p.371-381.
  • TÓTH T, KOVÁCS ZA & RÉKÁSI M. 2019. XRF-measured rubidium concentration is the best predictor variable for estimating the soil clay content and salinity of semi-humid soils in two catenas. Geoderma 342: 106-108.
  • VETTORI L. 1969. Boletim Técnico n.° 7.
  • WALKLEY A & BLACK IA. 1934. An examination of the degtjareff method for determining soil organic matter, and a proposed modification of the chromic acid titration method. Soil Sci 37: 29-38.
  • WAN M, HU W, QU M, LI W, ZHANG C, KANG J, HONG Y, CHEN Y & HUANG B. 2020. Rapid estimation of soil cation exchange capacity through sensor data fusion of portable XRF spectrometry and Vis-NIR spectroscopy. Geoderma 363: 114163.
  • WANG Y & WITTEN IH. 1997. Inducing Model Trees for Continuous Classes. Eur Conf Mach Learn 1-10.
  • WEINDORF DC, BAKR N & ZHU Y. 2014. Advances in Portable X-ray Fluorescence (PXRF) for Environmental, Pedological, and Agronomic Applications. In: Advances in Agronomy, Elsevier, p. 1-45.
  • WEINDORF DC & CHAKRABORTY S. 2020. Portable X-ray fluorescence spectrometry analysis of soils. Soil Sci Soc Am J 84: 1384-1392.
  • WOLD S, SJÖSTRÖM M & ERIKSSON L. 2001. PLS-regression: a basic tool of chemometrics. Chemom Intell Lab Syst 58: 109-130.
  • XU D, ZHAO R, LI S, CHEN S, JIANG Q, ZHOU L & SHI Z. 2019. Multi-sensor fusion for the determination of several soil properties in the Yangtze River Delta, China. Eur J Soil Sci 70: 162-173.
  • XU Y, DU A, WANG Z, ZHU W, LI C & WU L. 2020. Effects of different rotation periods of Eucalyptus plantations on soil physiochemical properties, enzyme activities, microbial biomass and microbial community structure and diversity. For Ecol Manage 456: 117683.
  • YESILONIS I, SZLAVECZ K, POUYAT R, WHIGHAM D & XIA L. 2016. Historical land use and stand age effects on forest soil properties in the Mid-Atlantic US. For Ecol Manage 370: 83-92.
  • ZHANG Y & HARTEMINK AE. 2020. Data fusion of vis-NIR and PXRF spectra to predict soil physical and chemical properties. Eur J Soil Sci 71: 316-333.
  • ZHAO D, WANG J, ZHAO X & TRIANTAFILIS J. 2022. Clay content mapping and uncertainty estimation using weighted model averaging. CATENA 209: 105791.
  • ZHU Y, WEINDORF DC & ZHANG W. 2011. Characterizing soils using a portable X-ray fluorescence spectrometer: 1. Soil texture. Geoderma 167-168: 167-177.

Publication Dates

  • Publication in this collection
    03 Oct 2025
  • Date of issue
    2025

History

  • Received
    19 Jan 2025
  • Accepted
    15 July 2025
location_on
Academia Brasileira de Ciências Rua Anfilófio de Carvalho, 29, 3º andar, 20030-060 Rio de Janeiro RJ Brasil, Tel: +55 (21) 2391-7901 - Rio de Janeiro - RJ - Brazil
E-mail: aabc@abc.org.br
rss_feed Acompañe los números de esta revista en su lector de RSS
Ir para arriba Notificar error