ORIGINAL RESEARCH article

Front. Sustain. Food Syst., 20 May 2026

Sec. Land, Livelihoods and Food Security

Volume 10 - 2026 | https://doi.org/10.3389/fsufs.2026.1748403

Analyzing techno-economic production efficiency with machine learning: the relationship between plant suitability and food security in Sub-Saharan Africa

  • 1. John von Neumann University, Kecskemét, Hungary

  • 2. Department of Economics and Management, J.Selye University, Komárno, Slovakia

  • 3. Department of Accounting and Auditing, Ferencz Rakoczi II, Transcarpathian Hungarian University, Beregovo, Ukraine

Abstract

Introduction:

In light of growing concerns about global food security, this study investigates how techno-economic indicators derived from production yield, market prices, and harvested area can be used to assess crop suitability in Sub-Saharan Africa.

Methods:

Using open-access statistical data from the Food and Agriculture Organization for the period 2000–2023, we apply machine learning (ML) models—Random Forest (RF) and multilayer perceptrons (MLP)—to predict a suitability proxy (Yield × log(Price)). In addition, we perform profitability-based clustering and trend analysis for five representative countries. In contrast to traditional ecological approaches, crop suitability is interpreted here in terms of economic viability. This perspective is particularly relevant in regions that are highly vulnerable to climate change.

Results:

The results indicate that yield is the most important predictor of both profitability and suitability, highlighting the critical role of agronomic performance in ensuring food security.

Discussion:

Our framework may also serve as a reference for evaluating high-tech agricultural solutions, such as IoT-based precision farming and remote sensing-based monitoring systems, thereby supporting evidencebased policy design in environments with limited technological capacity.

1 Introduction

Ensuring global food security has become an increasingly urgent challenge in the context of climate change, population growth, and resource constraints. In Sub-Saharan Africa, agricultural production systems are particularly vulnerable due to limited access to inputs, technological constraints, and high exposure to climatic variability. Identifying crops that are suitable under these conditions is therefore essential for improving food security and supporting evidence-based agricultural policies.

Traditionally, crop suitability has been assessed primarily from a biophysical or agroecological perspective, focusing on soil characteristics, climatic conditions, and crop-specific physiological requirements. While these approaches provide valuable insights into the biological potential of crops, they often overlook the economic dimension of agricultural decision-making. Farmers’ production choices are strongly influenced not only by ecological conditions but also by expected profitability. Consequently, crop suitability may also be interpreted from a techno-economic perspective, where economic viability becomes a key determinant of sustainable production.

This study examines the techno-economic interpretation of crop suitability by focusing on profitability estimated from crop yield and market prices. Instead of measuring suitability purely in a biological sense, we link it directly to the economic dimension of food security. In this framework, suitability is interpreted as an economic outcome emerging from realized yields and market conditions rather than as a theoretical biophysical potential.

Empirically, we apply machine learning methods—Random Forest (RF) and multilayer perceptrons (MLP)—together with cluster analysis to data obtained from the Food and Agriculture Organization (FAO). The analysis covers countries in Sub-Saharan Africa, as defined by the FAO classification, for the period 2000–2023. Using FAOSTAT data, we identify the most profitable crops and analyze their temporal dynamics across the region.

Our profitability estimates also serve as a low-tech reference for evaluating the potential benefits of precision agriculture technologies, such as IoT-based soil monitoring or data-driven crop management systems. The crops analyzed—including millet, beans, cassava, and others—exhibit different levels of technological sensitivity, allowing us to assess the relative effectiveness of precision agriculture interventions in a region where access to modern agricultural inputs remains limited.

In this study, biophysical constraints are treated as implicit conditions reflected in realized yields, while suitability is interpreted as an economic outcome rather than a biological potential. From a policy perspective, our results highlight that stabilizing agricultural production through targeted yield-enhancing interventions (e.g., input subsidies or improved agronomic practices) may have a greater impact on food security than policies focusing solely on managing price volatility.

2 Literature review

This literature review is narrative and exploratory in nature. Its objective is not to provide a systematic or bibliometric synthesis of existing studies, but to position the present research at the intersection of three related strands of literature: (i) crop suitability assessment, (ii) techno-economic interpretations of agricultural performance, and (iii) the application of machine learning methods in agricultural and food security research.

The reviewed studies were selected based on their relevance to yield-based productivity analysis, economic aspects of food security, and data-driven agricultural modelling. The purpose of the review is to illustrate how yield- and price-related outcome variables have been used in predictive and decision-support frameworks, and to motivate the proposed suitability proxy and modelling approach rather than to exhaustively map the field. In this context, ensuring the sustainability of global food systems while feeding a rapidly growing population has become one of the central challenges of the 21st century. In this context, increasing agricultural productivity without imposing additional environmental burdens has emerged as a key objective. A seminal global analysis by Foley et al. (2011) demonstrates that agricultural production can be intensified while limiting environmental pressures through improved land-use efficiency and yield improvements. Their work introduced the concept of sustainable intensification, which has since become a central paradigm in agricultural development research and forms an important conceptual foundation for the present study.

Empirical evidence suggests that Sub-Saharan Africa (SSA) faces particularly large productivity challenges. van Ztersum et al. (2016) show that the yield gap—the difference between potential and realized yields—is substantially larger in SSA than in most other regions of the world. Closing this gap is therefore widely considered a key pathway toward improving regional food security. At the same time, the development of commercial value chains remains a critical challenge in many input-constrained agricultural systems (Swinnen and Kuijpers, 2020). Technology adoption and knowledge diffusion play an important role in this process. For instance, Sseguya et al. (2021) demonstrate that demonstration plots can significantly accelerate the diffusion of improved agricultural technologies, including improved seeds, cereals, and tillage practices. However, technology adoption in SSA remains constrained by structural factors such as small farm sizes, limited access to credit, and time constraints faced by farmers (Maertens et al., 2020). Similarly, Mgendi et al. (2022) find that training and demonstration programs substantially increase the adoption intensity of new technologies, reporting an average increase of 64.4%. Nevertheless, the authors emphasize that demonstration initiatives cannot function effectively in isolation and must be complemented by extension services, input markets, and functioning value chains.

Technology transfer and knowledge exchange represent another important dimension of agricultural development in SSA. Mgendi et al. (2019) examine the impact of technology imports from China and Japan in Tanzania and Kenya and conclude that successful adoption depends not only on the technological characteristics themselves but also on their compatibility with local farming systems, infrastructure, and financing conditions. Their findings highlight the importance of the research–farmer–technology triad, involving collaboration between universities or research institutes, farmers, and technology providers. Similar conclusions are drawn in a World Bank report on technology adoption in SSA (Hörner et al., 2021), which shows that the diffusion of complex agricultural technologies—such as precision agriculture tools—depends heavily on farmers’ access to knowledge, demonstration programs, and research support. When research institutions actively participate in demonstration activities, the likelihood of technology adoption increases significantly.

Several studies also emphasize the structural constraints that characterize agricultural production systems in SSA. Jayne et al. (2014) describe African agriculture as predominantly small-scale and input-poor, facing persistent systemic challenges such as insecure land tenure systems and weak logistics infrastructure. Sheahan and Barrett (2017) further document the extremely low use of modern agricultural inputs in the region, attributing this pattern to high prices, limited availability, and insufficient knowledge among farmers. Earlier work by Nindi (1993) similarly stresses the importance of basic infrastructure, education, and input accessibility in enabling agricultural transformation. More recent research confirms that technological, structural, and organizational barriers continue to constrain agricultural development in the region (Nyaga et al., 2021).

Within this context, the adoption of precision agriculture (PA) technologies remains limited in most SSA countries. Nyaga et al. (2021) show that most existing research on precision agriculture in the region is concentrated in relatively technologically advanced countries such as South Africa, Nigeria, and Kenya. Furthermore, many pilot projects are conducted at the field or experimental scale rather than at the farm-system level. In more than 20 countries in the region, virtually no research on precision agriculture has been conducted. The authors emphasize that the predominance of smallholder farming systems significantly constrains the diffusion of such technologies. Bashiru et al. (2024) provide a comprehensive mapping of smart farming technologies in SSA and find that advanced technologies—including drones, IoT sensors, artificial intelligence, and GIS-based systems—remain relatively rare among smallholders. Key barriers include insufficient infrastructure, limited digital skills, financing constraints, and the lack of locally adaptable technological solutions. In addition to strengthening technology transfer and knowledge dissemination, the authors emphasize the importance of adapting innovations to the specific socio-economic context of smallholder agriculture.

More broadly, Klerkx and Rose (2020) highlight that agricultural technological transitions are inherently complex processes. Social acceptance, limited access to inputs, and insufficient adaptive capabilities frequently hinder the successful adoption of new technologies. Their review of precision agriculture tools used in Africa—including remote sensing systems, soil sensors, and mobile applications—shows that although current adoption rates remain relatively low, several promising pilot initiatives have emerged in countries such as Senegal and Ghana. Similarly, Jiang et al. (2016) analyze Chinese-funded Agricultural Technology Demonstration Centers (ATDCs) operating in Africa and identify several challenges related to the sustainability of these initiatives, including weak knowledge transfer mechanisms and limited integration into local agricultural systems. The broader relationship between research institutions, higher education, and agricultural communities is examined by Egeru et al. (2023), who emphasize the role of universities in supporting agricultural innovation through research farms, practical training programs, and collaborative knowledge exchange with farmers. Successful partnerships, according to their findings, require long-term commitment, sustainable financing, effective knowledge transfer mechanisms, and strong local adaptation.

Recent technological developments have also increased the importance of data-driven approaches in agriculture. Kamilaris et al. (2017) provide a comprehensive overview of the application of Big Data in agriculture and highlight the growing role of machine learning in agricultural prediction tasks. They emphasize that robust data preprocessing, model interpretability, and visualization are essential components of reliable predictive frameworks. Liakos et al. (2018) similarly review machine learning applications in agriculture and demonstrate that nonlinear models such as Random Forest and multilayer perceptrons (MLP) often outperform traditional statistical models in complex agronomic prediction tasks. Focusing specifically on African agriculture, Ly (2021) discusses the opportunities and challenges associated with machine learning applications in the region. While such approaches offer promising tools for yield prediction, soil monitoring, and risk management, their application remains constrained by limited data availability, infrastructural weaknesses, and the lack of locally adapted modeling frameworks.

Despite the extensive literature on crop suitability and agricultural technology adoption, several important research gaps remain. Most crop suitability assessments—such as those developed by the FAO or CGIAR—are primarily based on agroecological criteria, focusing on soil characteristics, climate conditions, heat accumulation, and precipitation patterns (Zabel et al., 2025; Ramirez-Villegas et al., 2013). While these models are valuable for identifying the biophysical potential of crops, they typically do not incorporate economic variables such as market prices or profitability. As a result, they do not fully capture the economic decision-making processes that influence farmers’ crop choices.

To address this limitation, the present study reinterprets crop suitability from a techno-economic perspective by integrating yield and price information into a profitability-based suitability indicator. Specifically, we employ a Yield × log(Price) proxy to capture the economic attractiveness of different crops. Rather than analyzing what could be produced under optimal biophysical conditions, the analysis focuses on what is economically viable under existing production conditions. This approach contributes to the relatively limited body of research on economic prediction models for input-constrained agricultural systems (Sheahan and Barrett, 2017; Mgendi et al., 2019; Mgendi et al., 2022).

Furthermore, existing machine learning applications in African agriculture are often limited to large-scale case studies (Nyaga et al., 2021), experimental farms, or field trials (Pokhariyal et al., 2023; Coman et al., 2025). Many studies also focus primarily on environmental indicators—such as NDVI or soil parameters—rather than integrating economic variables (Cedric et al., 2022; Tefera et al., 2025). In contrast, the present study introduces a hybrid techno-economic indicator that combines yield and price information to approximate profitability while also capturing the technological sensitivity of different crops.

A further contribution of this research lies in its broad geographical and temporal scope. The analysis covers 37 countries in Sub-Saharan Africa and multiple crop types over the period 2000–2023 using FAOSTAT data. This enables the identification of regional patterns and long-term trends in techno-economic crop suitability. In addition to predictive modeling using Random Forest and MLP algorithms, the study also applies clustering techniques to classify crops and countries according to their profitability dynamics. Consequently, the research not only develops predictive models but also provides a typological framework that can support strategic agricultural policy design in resource-constrained environments. Building on these strands of literature, the present study adopts a techno-economic perspective on crop suitability and applies machine learning methods not for algorithmic benchmarking, but to interpret the structural relationship between yield, price signals, and food security outcomes in Sub-Saharan Africa.

After reviewing the literature and assessing the research gap, we formulated the following hypotheses:

H1: Yield and revenue per hectare are the strongest predictors of crop suitability.

H2: Crop type is a relevant predictor of suitability in the techno-economic model.

H3: The clustering of countries is related to the diversity of their crop portfolios: countries with a higher cluster ratio typically exhibit a narrower crop structure.

These hypotheses are tested using the machine learning and clustering framework described in the following section.

3 Materials and methods

Guided by the conceptual framework outlined in the literature review, this section describes the data sources, the construction of the techno-economic suitability proxy, and the machine learning methods used to analyze its determinants across crops and countries in Sub-Saharan Africa.

3.1 Data used

We began our research by collecting data. We used publicly available data from FAO databases, ensuring transparency and adherence to data usage guidelines relevant to ethical research practices (FAO, 2023a, 2023b). Given that the target area of the research is Sub-Saharan Africa, we selected crop plants that are important from a research perspective.

Millet and sorghum are suitable for modeling climate adaptation due to their drought tolerance. Sweet potato and bean (cowpea) are key crops for nutritional security because of their calorie, vitamin, and protein content. The main advantage of cassava is its low input requirement, as it can grow without artificial fertilization or special cultivation practices; therefore, it is well suited for comparing precision and low-tech farming. The high market value of groundnut and rice ensures adequate income and livelihood for growers. As a local specialty, fonio can be considered a niche crop, as it also has export potential due to growing international demand.

For each crop, we collected data on area harvested (ha), yield (t/ha), production quantity (t), and producer price (USD/t) for the period 2000–2023 from 37 countries in the region. While FAOSTAT provides the most comprehensive and standardized cross-country agricultural dataset currently available, it should be noted that country-level aggregation may mask within-country heterogeneity in production conditions and technology adoption.

As with crops, scientific considerations were taken into account when selecting indicators. Area harvested characterizes food availability and system-level adaptation. An increase over time may indicate climate adaptation or market expansion. If the harvested area is large, the crop can be considered a staple from the perspective of food security. It also has policy relevance: if high-yield crops do not expand in cultivated area, this may indicate a technological deficit.

Yield (t/ha) serves as a proxy indicator of agronomic suitability and performance in this context, particularly when direct suitability metrics are not available. For a given crop, high yields suggest that the crop is well adapted to local conditions and that the soil–climate combination is suitable for cultivation, while local technologies function effectively. Therefore, this indicator can be used as a baseline variable for suitability score models. For example, it can serve as a dependent variable or as a component within Machine Learning (ML) models when examining the results of environmental impacts.

Production quantity is a useful indicator of output, resilience, and national food supply. Changes in production may indicate the effects of climate change or the introduction of new technologies, and can therefore be used in AI-driven predictive analyses (e.g., “what happens if...?” scenarios). Producer price reflects economic viability and food security from the farmer’s perspective (i.e., how much the crop is worth to the producer). High prices may reflect strong demand or limited supply, which can indicate food access challenges. Conversely, persistently low prices may signal producer losses and long-term economic unsustainability. This indicator can also be used in AI/ML models to predict future price dynamics under changing climatic conditions, thereby contributing to policy recommendations aimed at improving price stability and access.

We calculated additional variables based on the downloaded data (Table 1). Revenue per hectare (USD/ha) provides a useful basis for comparing the performance of different technologies and can also be used as a GIS layer to visualize where production may be economically viable. The suitability proxy models the suitability of individual crops. The machine learning models are used primarily for structural interpretation and feature contribution analysis rather than for forecasting an exogenous target variable.

Table 1

VariableFormulaWhat does it show?
Revenue/haYield × PriceIncome potential, profitability proxy
Total revenueProduction × PriceGDP-effect
Suitability proxyYield × log(Price)For complex ranking

Computed variables.

Source: authors’ own.

Our methodological approach in this area is aligned with the findings of Kamilaris et al. (2017), who demonstrate the applicability of machine learning techniques to yield- and market-related agricultural data for predictive and decision-support purposes. Differences in crop-specific error rates primarily reflect heterogeneity in production outcomes rather than systematic model bias caused by data imbalance. Machine learning methods are particularly suitable for this type of analysis because they can capture nonlinear relationships between agronomic and economic variables without requiring strong parametric assumptions.

Using the downloaded and computed variables, it is possible to examine which crops fit well within a given environment (crop suitability) and how secure the food supply (access, quantity, price) may be under the introduction of new technologies such as remote sensing, AI, IoT, and GIS.

Our sample is characterized using the prepared descriptive statistics. Descriptive statistics relevant to the research are presented here, while additional results can be found in the Supplementary Materials. First, we analyzed the yields of the examined crops (Figure 1).

Figure 1

The highest-yielding crops in Sub-Saharan Africa are sweet potatoes, with yields gradually increasing over the years, and rice, which shows relatively high but largely stagnant yields. The yield trends by country are shown in Figure 2.

Figure 2

The highest yields are observed in Cabo Verde (by far the highest) and Senegal, while Eritrea and Djibouti are at the lower end of the distribution. We also examined the total and per-hectare income generated from the crops (Table 1).

Sorghum and rice generate the highest total income, while sweet potatoes and rice produce the highest income per hectare (see Table 2).

Table 2

CropRevenue_per_ha(USD/ha)Total_revenue(USD)
Beans, dry76898,000,000
Sweet potatoes3,338119,000,000
Millet310148,000,000
Rice1,172182,000,000
Sorghum317204,000,000

Revenues from crops.

Source: authors’ own.

To assess crop suitability, we use observed yield (t/ha) as a proxy for local environmental compatibility. Areas with consistently high yields are interpreted as naturally suitable zones without high-tech intervention. Economic viability is estimated based on gross income per hectare (USD/ha), calculated by multiplying observed yield by the producer price. This serves as a benchmark for assessing the potential impact of precision farming technologies. Harvested area is used as a proxy for adaptation and systemic integration, modeling traditional dominance or technological expansion over time.

3.2 ML model building

The following software was used for data recording, data preparation, and modeling:

Excel for Mac v.16.89.1,

Jamovi 2.7.5.0,

Orange3-3.39.0-Python3.10.11.

The recorded data were prepared in the first step. As an initial step, we standardized the units of measurement and variable names, after which we imputed the missing data. Whenever possible, missing values were imputed based on the time trend. Data that could not be imputed in this way were replaced with the group average. Missing values were first imputed using crop-group means. Remaining missing values were then filled using the average/most-frequent method, where numerical variables were replaced by their mean and categorical variables by their most frequent category. Finally, rows containing missing producer price values were deleted.

Fields such as Country and Crop were converted into dummy vectors using one-hot encoding with the Continuize discrete variables – one feature per value procedure. Through standardization, variables were transformed to μ = 0 and σ2 = 1. Standardization supported faster convergence, particularly for parametric models such as logistic regression and neural networks, which are sensitive to input scale.

Data preparation was followed by the selection of the appropriate machine learning (ML) model. Our goal was to determine how effectively the crop suitability score can be predicted. We also examined which variables most strongly influence the cluster membership of a given observation (country–year–crop combination); therefore, we performed a feature importance analysis.

The Random Forest model was chosen as the initial model due to its robustness and practical advantages. It handles categorical variables efficiently, is relatively insensitive to data scaling (although standardization was applied regardless), and provides stable performance with minimal hyperparameter tuning. We used 100 trees to ensure prediction stability while maintaining computational efficiency. The depth of the trees was limited by setting the minimum number of samples per leaf to 5 in order to reduce overfitting. Training reproducibility was ensured by fixing the random seed, allowing the results to be consistently reproduced across runs.

Machine learning analyses were performed in Orange, where the training and test datasets were explicitly separated using the Test and Score widget. Model performance was recorded on both datasets, allowing direct comparison of the results. The program used 66% sampling for training, while the remaining 34% was used for testing. This procedure was complemented with multiple replication validation (10 repetitions) to ensure robustness. Performance metrics (MSE, MAE, R2) were displayed separately for the test data within the Orange interface.

The Random Forest served as a baseline model against which alternative algorithms could be compared to determine whether a better-performing model could be identified. However, the objective of the machine learning application in this study is explanatory insight and policy relevance rather than exhaustive algorithmic benchmarking.

Feature Importance obtained from the Random Forest model helps to understand how the variables used for clustering contribute to the target variable. This indicator is therefore not directly related to the clustering procedure itself, but was used to validate the role of the proxy variable. It indicates which factors determine this proxy, after which the clusters created along the proxy and other variables were examined. Feature importance also allows the interpretation of clusters along a techno-economic axis (for example, identifying clusters that are primarily yield-driven versus those that are more price-sensitive).

The Random Forest model was chosen primarily to support the objectives of this study. We examine how different techno-economic factors contribute to the economic suitability of crops. Accordingly, Random Forest was selected instead of Principal Component Analysis (PCA) to identify the most influential variables. Although PCA is an effective linear tool for identifying principal components, its orthogonal transformation does not provide a direct input–output relationship. Moreover, PCA components are composite variables rather than variable-specific indicators, which can make interpretation more difficult. In contrast, the Random Forest method can detect nonlinear relationships and directly quantify the contribution of each input variable to the prediction of the target variable (suitability proxy), thereby providing interpretable importance indicators.

This is particularly relevant because clustering algorithms incorporate variables into the distance metric but do not provide explicit variable weight or importance indicators. Therefore, Random Forest-based analysis was used as an additional interpretation tool.

Unfortunately, it was not possible to display the learning curve in the Orange version used for the study, as the Learning Curve widget required for this analysis was not available. Overfitting was therefore assessed in an alternative way based on train/test performance indicators and a scatter plot comparing predicted and true values (Figure 3). Although this is not a classical learning curve, it clearly illustrates how the model behaves on both training and test data. Points close to the 45° diagonal indicate accurate predictions.

Figure 3

The Random Forest MSE (0.223) and R2 (0.998) values indicate a very high level of explanatory power and prediction accuracy. This raises the possibility of data leakage. However, the main reason for this is that the target variable and some predictors are structurally linked. Furthermore, the seed was fixed (replicable training) to prevent random variation from influencing the results, and target coding was not used for categorical variables (e.g., country, year) whose predictive contribution was marginal. Consequently, the model does not perform temporal or spatial forecasting. Another possible explanation for the high R2 is that the dataset is relatively regular and structured, with few missing values and well-scaled numerical variables. In this context, this can be considered an advantage rather than a limitation. To further verify the robustness of the model, prediction accuracy was evaluated on separate training and test datasets, and consistent performance across the validation runs confirmed that the results are not driven by overfitting. The Random Forest model therefore serves primarily as an interpretable baseline for identifying structural relationships in the data rather than as a tool for optimizing predictive performance.

Following the modeling step, clusters were created using k-means cluster analysis based on the variables Revenue_per_ha, Yield, Sustainability_proxy, and Producer_price. Together, these variables represent different dimensions of agricultural performance and resilience. The objective was to identify distinct crop types within Sub-Saharan agriculture according to their techno-economic profiles.

Observations were segmented according to yield, income, revenue per hectare, and the calculated suitability value. The optimal number of clusters was determined using the Elbow method. According to this approach, a clear breakpoint appears at k = 3 clusters, which was therefore considered optimal.

These clusters represent different groups that may reflect distinct agricultural systems and production strategies. We subsequently compared countries across clusters and analyzed which system types dominate in each case. Based on these results, policy-relevant recommendations can be formulated for decision-makers.

The applied methodology is designed to ensure transparency and reproducibility while focusing on the structural interpretation of techno-economic suitability rather than on short-term forecasting accuracy or algorithmic performance benchmarking.

4 Results

The Random Forest model achieved a mean absolute error (MAE) of 0.143, which is low relative to the scale of the suitability proxy variable (range approximately 4.5–10.8; mean ≈ 7.3). This implies that the average prediction error represents less than 2% of the mean value of the target variable. In comparison, the MAE of the Neural Network model was 0.168, while Linear Regression produced a substantially higher MAE of 0.309 (Table 3). While the R2 value of 0.997 indicates that a large proportion of the variance is explained by the model, the low MAE confirms that the predictions are not only statistically accurate but also practically precise.

Table 3

ModelMSERMSEMAEMAPER2
Neural network0.1050.3250.1682.10592E + 140.999
Random forest0.4460.6680.1432.72359E + 130.997
Linear regression0.5240.7240.3092.84122E + 130.996
Tree1.1081.0530.2552.07764E + 130.992

ML models comparison.

Source: Authors’ own (from Orange output).

Based on the Random Forest results, yield (t/ha) proved to be by far the most important predictor of the suitability proxy (RF importance = 0.974), accounting for more than 97% of the total predictive contribution (Table 4). Yield also exhibits a very strong linear relationship with the suitability proxy (Yield × log(Producer price)), as reflected by the univariate regression coefficient (175051.6). This suggests that even a simple linear model using yield alone would explain most of the variation in the target variable. Although revenue per hectare is statistically significant in the univariate regression (4578), its relative importance in the Random Forest model is almost 50 times smaller than that of yield. This likely reflects strong correlations with other variables, particularly yield.

Table 4

ImportanceVariableUnivar. reg.RReliefRandom Forest (RF)
1Yield (t/ha)175,051.5460.1290.974
2Revenue_per_ha (USD/ha)4,577.7950.060.021
3Crop = Sweet potatoes2,005.15100
4Crop Group = Roots and tubers2,005.15100
5Crop Group = Cereals562.62200

Importance of examined variables in crop suitability.

Source: authors’ own (from Orange output).

The Regression ReliefF (RRelief) importance values provide additional insight into variable relevance. This method evaluates how changes in predictor values correspond to changes in the target variable within local neighborhoods of the data. Its advantage lies in its ability to detect nonlinear and local patterns and to handle categorical variables and interactions. However, it can be sensitive to noise and does not provide an explicit model structure. The RRelief results also indicate a relationship between yield and the target variable (RRelief = 0.129). Among the variables not included in Table 4, Year (RRelief = 0.126) and certain country categories (e.g., Mali and Senegal) also exhibit some local influence.

Country, Year, and Crop variables show negligible importance in the Random Forest model (RF importance = 0.000–0.001). The minimal contribution of country- and year-specific variables suggests that yield already captures many of the structural differences associated with geographic and temporal variation. The predictive contribution of Revenue_per_ha was also limited (RF importance = 0.021), likely due to multicollinearity with yield.

Based on these findings, hypothesis H1 can be considered partially supported. Yield explains the vast majority of the variation in the suitability proxy, while revenue per hectare does not provide additional predictive power once yield is included in the model. Given the negligible importance of crop type variables, hypothesis H2 is not supported.

The performance of the tested machine learning models is summarized in Table 3. According to the evaluation metrics (MSE, MAE, R2), the Neural Network achieved the highest explanatory power. However, the Random Forest model demonstrated similarly strong predictive performance while offering greater interpretability and stability. Linear Regression produced weaker results, indicating that purely linear relationships cannot adequately capture the structure of the data. The Decision Tree model was overly simple and less accurate. These findings suggest that nonlinear relationships and interactions play a key role in predicting techno-economic crop suitability. Similar conclusions were reported by Liakos et al. (2018), who emphasize the advantages of nonlinear machine learning models in agricultural prediction tasks. Given its strong predictive performance and interpretability, the Random Forest model was selected as the primary model for further analysis.

The predictive performance of the Random Forest model across crops is illustrated in Figure 3. The horizontal axis shows the observed values of the suitability proxy, while the vertical axis represents the values predicted by the model. The observations are closely aligned along the 45-degree diagonal, indicating a high level of prediction accuracy. This pattern is consistent with the high R2 and low MAE values reported above. The model performs well across most crops without evident systematic bias.

However, some differences emerge when prediction errors are examined at the crop level (Figure 4). Sweet potato shows a relatively higher prediction error (mean error: 0.23; standard deviation: 0.54), whereas predictions for crops such as millet, beans, and sorghum are particularly accurate. The higher frequency of sweet potato observations in the dataset may have influenced the learning process and contributed to this pattern.

Figure 4

The k-means cluster analysis identified three clearly distinguishable groups of crop production profiles based on the variables Yield, Revenue_per_ha, Producer_price, and Suitability_proxy.

Cluster 3 represents the dominant production pattern, accounting for approximately 72% of the observations. This cluster is characterized by relatively low yields and moderate prices, reflecting the typical techno-economic profile of staple crop production in Sub-Saharan Africa.

Cluster 1 represents high-performance crops with both high yields and high revenue levels. These crops may include, for example, sweet potato or rice in particularly productive years. Such crops often benefit from more intensive input use and favorable production conditions.

Cluster 2 includes crops with relatively lower yields but higher market prices. In these cases, profitability is driven primarily by price premiums rather than high agronomic performance. These crops may correspond to niche or market-oriented products such as beans or fonio.

These clusters provide a useful typology of crop production systems and may support the design of targeted agricultural interventions and crop rotation strategies, particularly in resource-constrained environments.

The country–cluster distribution highlights the regional structure of techno-economic crop production profiles (Figure 5). Cluster 3 dominates the dataset, representing approximately 81% of all observations across countries. This indicates that most crop production in Sub-Saharan Africa is based on low-input, resilient staple crops with relatively low profitability.

Figure 5

Nevertheless, several notable exceptions can be observed. In the case of Cabo Verde, all observations fall into Cluster 1, indicating exceptionally strong techno-economic performance compared with other countries in the dataset. The country’s island geography, relatively intensive land use, and higher technological inputs may contribute to these results. However, this pattern should be interpreted cautiously due to the limited number of observations.

Eritrea also shows a relatively high proportion of observations in Cluster 1. However, data availability for this country is limited and may contain estimated values, which could influence the clustering results. Therefore, the interpretation of Eritrea’s cluster position should be treated with caution.

Senegal represents another interesting case. Although most of its observations still belong to Cluster 3, approximately 9% fall into Cluster 1, which is relatively high compared with most other countries in the dataset. This may reflect regional differences in agricultural conditions, the presence of irrigated farming systems, and the cultivation of crops with higher market value.

In contrast, several countries—including Burkina Faso, Burundi, Togo, and Uganda—are almost entirely represented in Cluster 3. This suggests relatively limited agronomic performance and profitability across the observed crop types.

The cluster analysis also provides insights into the relationship between national crop portfolios and techno-economic performance. Countries with a higher proportion of observations in the top-performing cluster tend to specialize in a narrower set of high-performing crops, whereas countries with more diversified crop portfolios typically show lower overall productivity and profitability.

This pattern is consistent with the argument presented by Swinnen and Kuijpers (2020), according to which agricultural specialization is often associated with efficiency gains, while diversification strategies may primarily support resilience rather than profit maximization. In the dataset, most countries exhibit a broad crop portfolio but low performance, resulting in a large share of observations in Cluster 3. In contrast, countries with a stronger presence in Cluster 1 tend to rely on a smaller number of highly productive crops. Based on these findings, hypothesis H3 can be considered supported.

Overall, the results indicate that yield is the dominant determinant of techno-economic crop suitability, while price effects play a secondary role. The cluster analysis further reveals structural differences in agricultural production profiles across Sub-Saharan countries, highlighting the importance of productivity improvements for enhancing food security.

5 Discussion

This study reinterprets crop suitability from a techno-economic perspective. Nevertheless, it is important to relate these findings to traditional ecological approaches. Classical suitability models—such as FAO EcoCrop or CGIAR-based climate–soil models (e.g., Zabel et al., 2025; Ramirez-Villegas et al., 2013)—primarily assess where crops can be grown based on environmental conditions. These approaches typically focus on agro-ecological constraints and potential productivity, while market conditions and profitability factors are not explicitly considered.

In contrast, the present study addresses a complementary but distinct question: not “Where can crops be grown?” but rather “Where is it economically viable to grow them?” By combining yield and market prices into a composite suitability proxy (Yield × log(Price)), the model captures both agronomic performance and economic relevance. Yield reflects realized environmental adaptability, while price incorporates the economic incentives associated with production. This techno-economic interpretation is particularly relevant in regions with limited access to inputs and high market uncertainty.

Although the datasets and variables used in ecological suitability models differ substantially from those used in this study—for example, ecological models typically rely on GIS-based soil and climate layers—some indirect comparisons are possible. For instance, countries such as Cabo Verde and Senegal are often classified as climatically marginal in ecological suitability assessments. The relatively high techno-economic suitability observed in our results suggests that technological adaptation, irrigation, or market integration may partially compensate for environmental constraints. Consequently, the proposed approach does not aim to replace ecological suitability models, but rather to complement them by incorporating an economic dimension that reflects the real-world viability of agricultural production.

The machine learning results indicate that agronomic factors play a dominant role in determining techno-economic crop suitability in Sub-Saharan Africa. Yield (t/ha) was by far the strongest predictor in the Random Forest model, accounting for more than 97% of the predictive contribution. This finding is consistent with the results reported by van Ztersum et al. (2016), who highlight the central role of yield in determining agricultural performance and food security outcomes. In low-input farming systems, physical production output often has a stronger influence on economic viability than price fluctuations. Although revenue per hectare is theoretically a relevant indicator of profitability, its contribution to suitability prediction was limited once yield was included in the model.

The negligible importance of country and year variables suggests that the structural determinants of techno-economic suitability are relatively consistent across space and time. This finding may indicate that productivity-enhancing technologies and management practices could potentially be transferred across regions with similar agro-economic conditions.

The cluster analysis further revealed distinct techno-economic production profiles among crops. These clusters provide a useful typology of agricultural production systems and may support the development of crop rotation strategies and targeted policy interventions. Identifying economically viable crop groups is particularly important for strengthening resilience and sustainability in resource-constrained agricultural systems.

The cluster structure also allows the formulation of differentiated policy implications. Crops belonging to Cluster 1 represent high-performance production systems characterized by high yields and high economic returns. In such contexts, policy efforts may focus on knowledge dissemination, technological scaling, and the strengthening of innovation networks. Potential measures include the establishment of demonstration farms, regional knowledge-sharing programs, and increased collaboration between research institutions and local universities. As emphasized by Klerkx and Rose (2020), the diffusion of precision agriculture technologies depends not only on technological availability but also on institutional and organizational capacities.

Cluster 2 is characterized by crops that achieve relatively high market prices despite lower yields. For these systems, the primary development opportunity lies in improving market access and strengthening value chains. Policy measures may include the development of producer cooperatives, the expansion of digital market platforms, improved logistics infrastructure, and the introduction of market information systems. Swinnen and Kuijpers (2020) similarly emphasize that improved market integration can significantly enhance agricultural profitability in developing regions.

Cluster 3 represents the dominant production pattern in the dataset and is characterized by relatively low yields and limited profitability. In these systems, agricultural production often prioritizes subsistence and resilience rather than market-oriented profit maximization. This pattern corresponds to the structural constraints described by Sheahan and Barrett (2017), including limited access to inputs, infrastructure deficits, knowledge gaps, and financial constraints.

Improving productivity in these contexts requires targeted interventions aimed at strengthening basic agricultural capacities. Previous studies emphasize the importance of input access, technology transfer, and institutional support (Jayne et al., 2014). Practical measures may include seed and fertilizer support programs, the introduction of climate-resilient crop varieties, training in low-input sustainable farming methods, and the gradual expansion of irrigation infrastructure. These interventions aim to improve yield stability and long-term resilience.

Several studies also suggest that yield-enhancing interventions often have a stronger impact on food security than price-based policy measures alone (van Ztersum et al., 2016; Nyaga et al., 2021). In addition, the promotion of high-yielding but resource-efficient crops may contribute to sustainable intensification strategies (Foley et al., 2011). In the present dataset, millet represents one example of such a crop.

Finally, the effectiveness of these interventions depends strongly on institutional capacity, knowledge networks, and research collaboration. The integration of agricultural R&D with university and extension systems has been identified as a key driver of innovation diffusion (Jiang et al., 2016; Egeru et al., 2023). Accordingly, the policy implications presented in this study should be interpreted as cluster-specific strategic directions informed by techno-economic patterns and existing literature, rather than as direct policy prescriptions derived from a single indicator.

6 Conclusion

This study examined the techno-economic suitability and economic viability of major crops across Sub-Saharan Africa by combining yield-based indicators with market price data. By integrating these variables into a composite suitability proxy (Yield × log(Price)), the analysis provides a profitability-oriented interpretation of crop suitability. The results indicate that yield plays a dominant role in determining techno-economic suitability, while price effects have a secondary influence. Crops such as cassava and millet perform relatively well under low-input farming conditions due to their resilience to both environmental and economic constraints.

By linking agronomic performance with economic incentives, this study proposes a framework that complements traditional ecological suitability assessments. While ecological models focus primarily on environmental conditions that determine where crops can be grown, the techno-economic approach applied here addresses where crop production is economically viable. This perspective is particularly relevant in resource-constrained agricultural systems, where farmers’ decisions are strongly influenced by both productivity and market conditions.

The study makes several contributions to the literature. First, it introduces a techno-economic reinterpretation of crop suitability that integrates agronomic performance and market incentives. The proposed suitability proxy provides a hybrid indicator capturing both production outcomes and economic value. Second, the analysis covers a long-term dataset spanning 23 years and 37 countries, allowing the identification of regional patterns and structural trends in agricultural performance. Third, the application of machine learning models provides insight into the relative importance of different techno-economic factors, highlighting the dominant role of yield in explaining suitability outcomes. Finally, the cluster analysis identifies distinct production profiles across crops and countries, enabling the formulation of differentiated policy implications for agricultural development strategies.

The findings also have practical policy relevance. In low-input agricultural systems, interventions aimed at improving yield stability—such as access to improved seeds, fertilizer programs, irrigation development, and agricultural extension—may have a stronger impact on food security than price-based policies alone. At the same time, the identification of economically viable crop groups may support the development of targeted crop diversification strategies and resilience-oriented agricultural planning.

Several limitations of the present study should be acknowledged. The interpretation of variable importance in Random Forest models is inherently limited, particularly when predictors are highly correlated. Methods such as SHAP values or partial dependence plots could provide a more detailed understanding of the local and global effects of predictors. However, the primary objective of this study was exploratory and explanatory rather than causal inference. Accordingly, Random Forest importance measures were used as an initial tool for identifying structural relationships in the data. Furthermore, the suitability proxy captures long-term techno-economic patterns rather than short-term market fluctuations or policy shocks. The results should therefore be interpreted as indicators of structural production conditions rather than precise forecasts of agricultural profitability. Future research could extend this framework by integrating additional variables such as climate variability, input intensity, or farm-level management practices. Combining techno-economic indicators with spatially explicit ecological models may also provide a more comprehensive understanding of agricultural sustainability in Sub-Saharan Africa. Overall, the findings suggest that improving yield stability through targeted technological and institutional interventions may represent one of the most effective pathways for strengthening food security in low-input agricultural systems.

Statements

Data availability statement

The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.

Author contributions

JK: Methodology, Supervision, Conceptualization, Software, Writing – review & editing. ZS: Writing – original draft, Writing – review & editing, Supervision, Conceptualization. DS: Visualization, Conceptualization, Investigation, Methodology, Writing – review & editing, Supervision. RB: Formal analysis, Conceptualization, Writing – review & editing, Software, Writing – original draft, Investigation. BK: Writing – original draft, Resources, Investigation, Formal analysis, Conceptualization, Methodology, Writing – review & editing.

Funding

The author(s) declared that financial support was not received for this work and/or its publication.

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declared that Generative AI was not used in the creation of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Correction note

This article has been corrected with minor changes. These changes do not impact the scientific content of the article.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

References

  • 1

    BashiruM.OuedraogoM.OuedraogoA.LäderachP. (2024). Smart farming technologies for sustainable agriculture: a review of the promotion and adoption strategies by smallholders in sub-Saharan Africa. Sustainability16:4817. doi: 10.3390/su16114817

  • 2

    CedricL. S.AdoniW. Y. H.AworkaR.ZoueuJ. T.MutomboF. K.KrichenM.et al. (2022). Crops yield prediction based on machine learning models: case of west African countries. Smart Agric. Technol.2:100049. doi: 10.1016/j.atech.2022.100049

  • 3

    ComanC.ComanE.GherheșV.BucsA.RadD. (2025). Application of remote sensing and machine learning in sustainable agriculture. Sustainability17:5601. doi: 10.3390/su17125601

  • 4

    EgeruA.LindowM.Muir LerescheK. (2023). University Engagement with Farming Communities in Africa. Routledge.

  • 5

    FAO. (2023a). FAOSTAT: crops and livestock products. Food and agriculture Organization of the United Nations Available online at: https://www.fao.org/faostat/en/#data/QCL (Accessed January 15, 2025).

  • 6

    FAO (2023b). FAOSTAT: producer prices food and agriculture Organization of the United Nations Available online at: https://www.fao.org/faostat/en/#data/PP (Accessed November 11, 2025).

  • 7

    FoleyJ. A.RamankuttyN.BraumanK. A.CassidyE. S.GerberJ. S.JohnstonM.et al. (2011). Solutions for a cultivated planet. Nature478, 337342. doi: 10.1038/nature10452

  • 8

    HörnerD.Perez-ArceF.ArslanA.BlackR.. (2021). Knowledge and Adoption of Complex Agricultural Technologies: Evidence from Sub-Saharan Africa. World Bank Policy Research Working Paper, (14b7585b-c58f-5a50-b3f4-475e308eb0ea). Available online at: https://hdl.handle.net/10986/36552 (Accessed August 1, 2024).

  • 9

    JayneT. S.ChamberlinJ.HeadeyD. D. (2014). Land pressures, the evolution of farming systems, and development strategies in Africa: a synthesis. Food Policy48, 117. doi: 10.1016/j.foodpol.2014.05.014

  • 10

    JiangL.HardingA.AnseeuwW.AldenC. (2016). Chinese agriculture technology demonstration centres in southern Africa: the new business of development. The Public Sphere8:736.

  • 11

    KamilarisA.KartakoullisA.Prenafeta-BoldúF. X. (2017). A review on the practice of big data analysis in agriculture. Comput. Electron. Agric.143, 2337. doi: 10.1016/j.compag.2017.09.037

  • 12

    KlerkxL.RoseD. (2020). Dealing with the game-changing technologies of agriculture 4.0: how do we manage diversity and responsibility in food system transition pathways?Glob. Food Secur.24:100347. doi: 10.1016/j.gfs.2019.100347

  • 13

    LiakosK. G.BusatoP.MoshouD.PearsonS.BochtisD. (2018). Machine learning in agriculture: a review. Sensors18:2674. doi: 10.3390/s18082674,

  • 14

    LyR.. (2021). Machine learning challenges and opportunities in the African agricultural sector – a general perspective. Available online at: https://farm-d.org/wp-content/uploads/2024/03/2107.05101.pdf (Accessed October 21, 2025).

  • 15

    MaertensA.MhangoW.MichelsonH.. (2020). The effect of demonstration plots and the warehouse receipt system on integrated soil fertility management adoption, yield and income of smallholder farmers: a study from Malawi’s anchor farms 3ie impact evaluation report 122 International initiative for impact evaluation. doi: 10.23846/TW4IE122

  • 16

    MgendiB. G.MaoS.QiaoF. (2022). Does agricultural training and demonstration matter in technology adoption? The empirical evidence from small rice farmers in Tanzania. Technol. Soc.70:102024. doi: 10.1016/j.techsoc.2022.102024

  • 17

    MgendiG.ShipingM.XiangC. (2019). A review of agricultural technology transfer in Africa: lessons from Japan and China case projects in Tanzania and Kenya. Sustainability11:6598. doi: 10.3390/su11236598

  • 18

    NindiB. (1993). Agricultural transformation in sub-Saharan Africa: the search for viable options. Nord. J. Afr. Stud.2, 142158.

  • 19

    NyagaJ. M.OnyangoC. M.WetterlindJ.SöderströmM. (2021). Precision agriculture research in sub Saharan Africa countries: a systematic map. Precis. Agric.22, 11741194. doi: 10.1007/s11119-020-09780-w

  • 20

    PokhariyalS.PatelN. R.GovindA. (2023). Machine learning-driven remote sensing applications for agriculture in India—a systematic review. Agronomy13:2302. doi: 10.3390/agronomy13092302

  • 21

    Ramirez-VillegasJ.JarvisA.LäderachP. (2013). Empirical approaches for assessing impacts of climate change on agriculture: the EcoCrop model and a case study with grain sorghum. Agric. For. Meteorol.170, 6778. doi: 10.1016/j.agrformet.2011.09.005

  • 22

    SheahanM.BarrettC. B. (2017). Ten striking facts about agricultural input use in sub-Saharan Africa. Food Policy67, 1225. doi: 10.1016/j.foodpol.2016.09.010,

  • 23

    SseguyaH.RobinsonD. S.MwangoH. R.FlockJ. A.MandaJ.AbedR.et al. (2021). The impact of demonstration plots on improved agricultural input purchase in Tanzania: implications for policy and practice. PLoS One16:e0243896. doi: 10.1371/journal.pone.0243896,

  • 24

    SwinnenJ.KuijpersR.. (2020). Inclusive value chains to accelerate poverty reduction in Africa. Jobs working paper, no. 37; World Bank. Available online at: http://hdl.handle.net/10986/33397 (Accessed November 14, 2025).

  • 25

    TeferaM. L.ZelekeE. B.PirastruM.MelesseA. M.SeddaiuG.AwadaH. (2025). Satellite-based machine learning for soil moisture prediction and land conservation practice assessment in west African drylands. Remote Sens.17:3651. doi: 10.3390/rs17213651

  • 26

    van ZtersumM. K.van BusselL. G.WolfJ.GrassiniP.van WartJ.GuilpartN.et al. (2016). Can sub-Saharan Africa feed itself?Proc. Natl. Acad. Sci.113, 1496414969. doi: 10.1073/pnas.1610359113,

  • 27

    ZabelF.KnüttelM.PoschlodB. (2025). CropSuite v1.0 – a comprehensive open-source crop suitability model considering climate variability for climate impact assessment. Geosci. Model Dev.18, 10671087. doi: 10.5194/gmd-18-1067-2025

Summary

Keywords

Africa, agriculture, crop, suitability, technology

Citation

Kárpáti J, Szeiner Z, Szabó D, Bacsó R and Kálmán BG (2026) Analyzing techno-economic production efficiency with machine learning: the relationship between plant suitability and food security in Sub-Saharan Africa. Front. Sustain. Food Syst. 10:1748403. doi: 10.3389/fsufs.2026.1748403

Received

17 November 2025

Revised

13 March 2026

Accepted

06 April 2026

Published

20 May 2026

Corrected

02 September 2026

Volume

10 - 2026

Edited by

Mohamed Shokr, Tanta University, Egypt

Reviewed by

Joshuva Arockia Dhanraj, Dayananda Sagar University, India

Nayeli Montalvo-Romero, Misantla Higher Technological Institute (ITSM), Mexico

Updates

Copyright

*Correspondence: Zsuzsanna Szeiner,

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics