Abstract
This study aims to maximize the informational return of mobile environmental data acquisition systems operating on road networks. We propose a two-step survey design framework that fuses spatially dense auxiliary data describing landscape heterogeneity with road network information and formulates route planning as a combinatorial optimization problem. An information-rich initial route is constructed using fuzzy clustering and orienteering optimization, followed by an economization step that preserves information content through convex hull analysis in auxiliary data space. The framework is evaluated using a large-scale (4,500 km2) case study of mobile Cosmic Ray Neutron Sensing (CRNS) soil moisture measurements in central Germany. The optimized route yields more representative spatial coverage and reduced extrapolation requirements when regionalizing sparse CRNS measurements into gravimetric soil moisture maps, thereby lowering ontological uncertainty in regression-based soil moisture mapping compared to an empirically designed route of similar length. The results demonstrate that explicitly information-driven route design substantially improves survey efficiency and data quality. The proposed framework offers a general and transferable alternative to convenience-based mobile sampling strategies in Earth and environmental sciences.
1 Introduction
Measurement systems record data, understood here as coded units of information. Coding precision is generally imperfect (Taylor, 1982), and defines the data’s aleatory uncertainty (), which limits reliable information retrieval from the data. A system may collect vast data sets, but if they contain insufficient information, robust interpretations remain impossible. Therefore, data acquisition strategies must be designed to maximize the informational return of data sets relative to measurement costs (; ).
In the past decade, mobile Cosmic Ray Neutron Sensing (CRNS) devices mounted on cars or trains have become increasingly used for sparsely measuring soil moisture, a crucial state variable for understanding hydrological, energy, ecological, and climate processes on regional to global scales (; Williams et al., 2012; van Westen et al., 2008). Typical mobile CRNS campaigns cover areas from a few to several thousand square kilometers (; ; ; ; ; ; ; ; ; ). Because full spatial coverage of large regions is economically infeasible (), the vehicle’s route becomes critical. It must balance informational value and acquisition cost to optimize information return (). Designing such a cost-efficient route with high information return constitutes a combinatorial optimization task, commonly formulated as an Orienteering Problem (OP) (), where the vehicle traverses a directed road network to collect maximum informational profit within constraints such as daily travel distance and time, while adhering to regular traffic conditions.
As in many other disciplines, measurement strategies in the Earth sciences often follow convenience or purpose sampling approaches (), rooted more in empirical or traditional routines than in optimizing information return. highlighted the importance of optimizing geophysical survey design, inspiring subsequent developments of optimized strategies for tomographic experiments (; Stummer et al., 2004; Wilkinson et al., 2012). Beyond geophysics, defining optimal sparse sampling schemes is a major topic in Earth and environmental sciences, including seismological network (), stream sediment sampling (), and validation of remote sensing products (). Various disciplines, such as spatial statistics and soil science, have proposed and evaluated optimized sampling schemes (; ; ; Žížala et al., 2024). For instance, developed an optimal sparse sampling method for stationary soil moisture monitoring based on fuzzy cluster analysis, dividing the survey area into landscape units that capture its spatial heterogeneity and ensure representative sampling. Similar concepts have been explored in transportation networks, where Yang and Yang (2023) and Yang et al. (2024) developed information-based approaches for optimizing sensor placement under resource constraints. While focusing on static sensors, these studies share the objective of maximizing information return in road-network environments. For moving instruments sampling the spatial patterns of a landscape, found that mobile observers relocating to regions with higher spatial variability achieved significant improvements in field reconstruction compared to static observation networks. Linking an OP with a road network, formulate a profitable close-enough arc routing problem where profits are associated with roads and solve it heuristically. Wang et al. (2024) further addressed standardized sampling layouts integrated with optimal path planning aiming to evaluate regional economic development in China using OpenStreetMap data, nighttime light data, and previous statistical data. However, recognized the difficulty of measuring the information value of large and complex data sets. Entropy-based approaches are used widely for combining different types of data (; ; ). For example, used an entropy measure to optimally integrate previous information and spatial data in order to select new sites to set up precipitation monitoring network stations.
Striving to overcome currently practiced convenience sampling strategies in mobile CRNS surveys, we propose a generalized optimization framework for routing mobile sensing systems that is information return driven. This framework takes the approach of as a starting point, and combines Fuzzy C-Means (FCM) clustering for detecting land units as basis for an optimal sparse sampling with a routing problem on road network data. An informationally rich route using ant colony optimization (ACO) () is constructed. In a subsequent economization phase, we used convex hull-based filtering to get a distance-optimized route ensuring high information return of a mobile sensing campaign. We demonstrated its effectiveness in a soil moisture monitoring application, where a CRNS terrestrial vehicle system () is used for acquiring data.
We used an empirically constructed route following the currently practiced mobile CRNS survey design as a comparison case for our optimized route. have shown that such survey designs risk producing data sets with highly variable information content and poor spatial coverage, thereby limiting the reliability of soil moisture maps derived from sparsely measured CRNS data across the survey area.
Using different non-parametric regression solvers such as Random Forest (RF) and Artificial Neural Network (ANN) for soil moisture regionalization, we were able to prove that our optimized route enables the generation of soil moisture maps with lower uncertainty compared to those resulting from the empirical route. This allows us to state that using an information driven routing scheme is essential for robust predictions, such as soil moisture maps of the entire survey area, when information cost is also considered in the problem enforcing generally sparse sampling.
In the following sections, we will explain our generalized framework and apply it to acquire a CRNS data set in a 4,500 km2 area southwest of the city of Magdeburg, Germany. We will discuss the gains and risks of our approach and evaluate its performance against empirical sampling, by quantifying the uncertainty of soil moisture map generation with regard to the information content of the acquired data sets.
2 Methodology
2.1 General workflow
The methodology of this study can be described as a data fusion and optimization problem. Its major steps are summarized in the flowchart in Figure 1. Our method takes multiple auxiliary spatial data sets, e.g., mapped soil properties, as input. To assess spatial heterogeneity in the survey area, we construct features from these data that are suitable for a fuzzy c-means cluster analysis. This results in a division of the survey area into land units (clusters) that capture its spatial heterogeneity as described by the auxiliary data. Additionally, OSM road network data of the area are acquired as a second input type and filtered to exclude undesired path types, such as cycling lanes. By linking the fuzzy clustered map with the filtered road network, we identify which map grid nodes are reachable via the considered road network. Based on the fuzzy information we calculate an informational value for each road segment.
FIGURE 1
We construct an initial optimal route by heuristically solving an OP using ant colony optimization (ACO), ensuring that each cluster type is visited at least once. This initial route is informationally rich but may still be costly in terms of distance or time. To further economize the route, we select those road segments that optimally cover the volume of the convex hull spanned in the auxiliary data space, considering only the map grid nodes reached by the initial route. Using these segments, we construct a final route that preserves most of the information of the initial route while being economically optimized, e.g., with respect to distance. This route ensures that, in a subsequent map generation procedure regionalizing CRNS data into a spatially dense map by multiple regression with the auxiliary data, extrapolation requirements are minimized.
2.2 Feature engineering and fuzzy clustering
The t collocated auxiliary data sets, i.e., soil maps covering the survey area, were vectorized and merged into the n × t matrix X with n being the number of map grid nodes. We apply a Singular Value Decomposition (SVD) X = USVT. Here, S is a diagonal matrix containing the singular values and U and V contain the left and right singular vectors, respectively. We compute uncorrelated principal components in the t-dimensional auxiliary data space by multiplying U and S, resulting in an n × t matrix D. Its columns are uncorrelated and their amplitudes are scaled according to the corresponding variance, with the highest variance contained in the first column, which represents the dominating common spatial variability of all auxiliary data sets.
We subject D to a fuzzy c-means (FCM) clustering that iteratively minimizes the objective function J with respect to the k × n fuzzy membership matrix M and the k × t cluster center matrix C:
In Equation 1, dj is the jth row (sample) of D, and ci is the position of the ith cluster center (the ith row of C). The number of clusters is k, and f is the fuzzification exponent controlling the fuzziness of the system. Following , we choose f = 2. The membership matrix M, with elements mij, describes the spatial heterogeneity of all input data with respect to the cluster centers and is constrained as seen in Equation 2.
We compute a crisp cluster index vector h by applying core selection defuzzification (Van Leekwijck and Kerre, 1999)
Following , we also compute a vector q quantifying the trustworthiness of the crisp cluster assignment to every map grid node
We empirically set p = 1. While h and q are only used for visualization purposes, M provides fuzzy membership values that enable the definition of scores for informational value computations.
We repeatedly apply clustering for all k ∈ [2, 50]. To select an optimal k, we compute a relative residual error ε(rel) for each clustering solution that quantifies the information loss resulting from representing the elements djv ∈ D by their cluster centers and normalize it with the residual error for k=2, using Equation 5.
2.3 Initial route computation
The first step in constructing an initial route is the combination of clustered auxiliary data and road network data. The road network data are sourced from OSM, an open-access database (https://www.openstreetmap.org, accessed 20 March 2025) maintained by volunteers who collect and edit road information. As a community-driven database, OSM may contain inconsistencies due to varying contributor experience, and database inhomogeneities may exist because not all attributes are provided for each road segment. Since we cannot control data consistency from our user position, we assume that the provided data are correct. Available road network data include roads with different properties such as name, road type, number of lanes, and speed limits. Each road is divided into segments, which are uninterrupted sections indicated by start and end coordinates. Segments inherit the properties of the road to which they belong.
We filter the road network according to road type. Roads not accessible to vehicles, e.g., pedestrian paths or cycling lanes, are excluded from the considered network. Similarly, roads marked as non-public are removed. If no accessibility information is provided, roads are assumed to be publicly accessible. Speed limit information is acquired where available. For roads without explicit limits, we assume values based on road type and national traffic regulations. For OSM track-type roads, speed limits are additionally adjusted according to the track grade classification (grades 1–5), reflecting decreasing road quality. One-way roads are identified and if this property is missing, roads are assumed to be accessible in both directions. Segments that are disconnected from the rest of the network are removed to reduce computational load.
The filtered road network is then fused with the clustered auxiliary data to determine which map grid nodes are reachable by the vehicle. A CRNS system has a circular sensing radius of approximately 150 m (). Any grid node with its center being within one-third of this radius from a road segment is considered reachable. Each road segment is evaluated individually, and at the end of the process, each reachable grid node is linked to one or more road segments sufficiently close to it.
Next, we quantify the informational value gained by visiting a road segment. For each segment r, we compute a cluster membership vector by averaging the fuzzy membership vectors of all grid nodes covered by that segment. The Shannon entropy er of segment r is then calculated following Singh (2013)
In Equation 6, is the membership value of segment r with respect to the ith cluster. Lower entropy indicates a more certain assignment to a distinct cluster. The information value vr of a segment is defined in Equation 7 as
Applying a complement transformation, we convert the information values into segment weights
These weights serve as costs for selecting road segments during initial route construction, where a shortest path algorithm is applied. Using Dijkstra’s algorithm (), we identify the lowest-weighted segment from each cluster and compute the most information-rich pathways connecting them. The proposed metric intentionally favors road segments exhibiting strong affiliation to a particular cluster type. The objective is not to identify locations characterized by maximal uncertainty or transition behavior, but rather to identify representative sampling locations for each land unit (cluster). Transition zones are already represented through the fuzzy memberships () and subsequently through the convex hull-based economization procedure. Therefore, the information-value metric is designed to prioritize prototypical representatives of each cluster rather than regions of maximum cluster ambiguity.
Finally, we determine the optimal sequence in which these low-weighted segments should be visited. This ordering constitutes an OP formulated as a combinatorial optimization task. We solve it heuristically using an ACO algorithm, yielding an initial route that is information rich and optimal with respect to the cluster visitation sequence. Because ACO contains stochastic components, multiple independent runs were performed using random seeds to verify solution stability. For the presented application, the final route corresponds to the best solution obtained from the optimization procedure. However, actual driving costs may still be high, as route length optimization has not yet been addressed.
2.4 Economization
While the initial route is information rich, it is often unnecessary to acquire CRNS readings at every map grid node along it. In the economization step we identify those grid nodes that are critical for preserving the information content. We begin by constructing a convex hull () in the auxiliary data space using the grid nodes covered by the initial route. To illustrate the concept, Figure 2 shows a toy example in a 2D auxiliary data space defined by predictors A and B. Crosses represent all collocated samples (map grid nodes), while circles indicate the nodes on the initial route where CRNS measurements could be taken. Predictors A and B are normalized, so that all samples on the initial route fall in the interval [0, 1], yielding a unit box shown in Figure 2b. Red crosses lie outside this box and represent nodes whose auxiliary data states fall outside the range sampled along the route. Blue crosses indicate nodes not reachable by the initial route, but within the unit box represent nodes whose auxiliary data states are well represented by the route samples (encircled crosses).
FIGURE 2
A more refined differentiation can be obtained when computing the convex hull of the route samples (green polygon in Figure 2c). Only a subset of samples on the route defines the hull. Blue crosses indicate samples that can be interpolated from the hull vertices, whereas red crosses correspond to nodes requiring extrapolation. The orange and magenta lines indicate normalized distances from hull vertices used during interpolation/extrapolation, i.e., in a multiple regression process regionalizing CRNS measurements acquired for grid nodes defining the hull vertices.
Computing the convex hull for all samples reached by the initial route allows us to identify grid nodes that are essential for preserving the extrapolation fraction when regionalizing CRNS data by multiple regression using the auxiliary data as predictors. Since not all points on the route define the convex hull, only for the subset of hull vertices CRNS measurements are needed to avoid increasing the extrapolation fraction or distances. If the number of hull vertices becomes large, we identify an optimal subset of a predefined size that best preserves the hull volume.
To select this subset, we use a Particle Swarm Optimization (PSO) () algorithm, executed multiple times with different random seeds to enhance robustness. The best solution among all runs is retained for subsequent analyses. The result is a set of grid nodes that preserves the convex hull volume optimally while keeping computation time manageable. Importantly, the convex hull approach is essential for ensuring low extrapolation components during CRNS regionalization and for achieving representative spatial coverage. However, convex hull computation becomes increasingly expensive with growing dimensionality of the auxiliary space and sample size. For our study area (approximately 4,500 km2), computing the hull of all grid nodes reachable by the road network was computationally infeasible. Thus, reducing the problem’s dimensionality by first generating an initial route based on fuzzy-cluster-derived soil landscape descriptors is essential for tractability.
To construct the final route, we apply a modified version of the initial route framework. We first identify the road segments covering all grid nodes corresponding to the selected convex hull vertices. To connect these segments, we compute shortest pathways between the selected segments using A* algorithm (), this time focusing specifically on minimizing total route length. The final route is again constructed using ACO and represents an economized version of the initial route. This optimized route maintains high information return while minimizing driving cost.
2.5 Regression-based evaluation of information return
To evaluate the impact of survey design on the information content of the acquired data sets, we regionalize sparse CRNS soil moisture measurements into spatially continuous soil moisture maps using two independent non-parametric regression approaches: Random Forest (RF) regression () and feed-forward Artificial Neural Networks (ANNs) (). Both methods are capable of representing nonlinear relationships between soil moisture and environmental predictor variables while relying on fundamentally different model structures.
For both the empirical and optimized routes, gravimetric soil moisture derived from the CRNS measurements serves as the response variable y. Using X as predictor matrix, the regression problem can be written asIn Equation 9, f denotes the unknown nonlinear mapping between the auxiliary variables and soil moisture, and e aggregates uncertainties related to the predictor variables, the response variable, and the regression model itself. The function f is learned from spatially collocated instances in y and X and is subsequently used to predict soil moisture at locations covered by X for which no measurements are available in y. CRNS measurements assigned to the same map grid node are averaged prior to analysis. Data-related uncertainties of the individual CRNS observations and the predictor variables are not explicitly considered.
For both routes, 70% of the available samples are used for model training, while the remaining samples are used for validation and testing. Based on empirical sensitivity analyses, the RF model is implemented using 100 decision trees and a minimum leaf size of five samples. The ANN consists of a single fully connected hidden layer with 10 neurons using a hyperbolic tangent activation function and a linear output layer. Network parameters are optimized using the Levenberg–Marquardt algorithm. Model training is terminated when a mean squared error (MSE) of 1e-3 is reached. Model performance is evaluated using the MSE.
The purpose of employing two different regression approaches is not to identify the best predictive model, but to evaluate the sensitivity of regionalized soil moisture maps to model structure. Differences between RF- and ANN-derived soil moisture maps are therefore used as a comparative indicator of model-structure-dependent (ontological) uncertainty associated with the sparse measurement data sets and their information value. If a survey design provides representative coverage of the auxiliary-data space and requires little extrapolation, both regression methods are expected to converge towards similar spatial predictions. Conversely, larger differences between RF and ANN predictions indicate greater freedom in model extrapolation and therefore increased model-structure-dependent (ontological) uncertainty.
For reproducibility and ease of reference, the principal implementation parameters used throughout the proposed workflow are summarized in Table 1.
TABLE 1
| Component | Parameter | Value |
|---|---|---|
| SVD | Retained components | 6 |
| FCM | Cluster range | 2–50 |
| FCM | Fuzzification exponent f | 2 |
| ACO | Number of ants | 50 |
| ACO | Iterations | 500 |
| ACO | ρ - pheromone evaporation | 0.5 |
| ACO | α - alpha | 0.5 |
| ACO | β – beta | 3 |
| PSO | Swarm size | 500 |
| PSO | Maximum iterations | 150 |
| PSO | Topology | Global best |
| PSO | Repeated runs | 5 |
| RF | Trees | 100 |
| RF | Minimum leaf size | 5 |
| ANN | Hidden neurons | 10 |
| ANN | Hidden activation | Tanh |
| ANN | Output activation | Linear |
| ANN | Optimizer | Levenberg-marquardt |
| ANN and RF | Target MSE | 10–3 |
Main implementation parameters used throughout the proposed workflow.
3 Application of methodology
3.1 Survey area and database
The survey area is located southwest of the city of Magdeburg in Germany. It covers approximately 4,500 km2 and comprises a variety of landforms, including the Harz Mountains in the south, forested areas, and the Grosses Bruch, a wetland strip formed by a glacial valley and characterized by lowland meadows and numerous ditches. It also includes the Magdeburger Börde, a fertile plain that forms part of a loess belt extending along the south-eastern rim of the North German Plain. The area is densely populated and intensively used for farming (Figure 3). It covers more than 90% of the Bode catchment, which has been intensively studied over the past decade within the TERENO program (Zacharias et al., 2024). The area also includes several smaller test sites, such as Hohes Holz, the Selke catchment, Schäfertal, and the Holtemme (Wollschläger et al., 2016), that have likewise been the focus of detailed investigations within the Bode catchment.
FIGURE 3
To demonstrate our methodology, we deliberately restrict the auxiliary data to spatial data sets that are commonly used in comparable soil moisture estimation studies (). These include physical and chemical soil properties as well as basic topographic attributes. While this selection is not intended to be exhaustive, for example, precipitation data of the recent past could also be considered, it serves to illustrate the proposed workflow rather than to debate the optimal choice of predictors.
The considered maps are shown in Figure 4 and comprise soil bulk density, clay content, sand fraction, and soil organic carbon, all obtained from www.soilgrids.org (accessed 15 October 2024), as well as a digital elevation model with 200 m spatial resolution provided by the German Federal Agency for Cartography and Geodesy (BKG; https://gdz.bkg.bund.de/index.php/default/digitales-gelandemodell-gitterweite-200-m-dgm200.html, accessed 15 October 2024). The elevation data were resampled to the grid of the soil maps using 250 m node spacing, and a corresponding slope map was derived. Together, these maps form the spatially dense input data matrix X, which is subjected to feature engineering via singular value decomposition to obtain principal components.
FIGURE 4
Clustering these principal components allows us to describe the spatial heterogeneity of the survey area in numerical form through fuzzy cluster memberships and to identify clusters, representing soil-landscape descriptors following . By evaluating the relative residual error ε(rel) for each clustering solution (Figure 5a), we select 15 clusters as a suitable compromise for capturing the spatial heterogeneity of the survey area as described by the input maps (Figure 4). Increasing the number of clusters would only marginally reduce information loss due to grouping samples into more clusters, whereas solutions with fewer clusters show a rapid increase in information loss, as indicated by the residual error. Note, the residual error decreases monotonically with increasing numbers of clusters and reaches zero for n = k. Therefore, the optimal solution is not identified by the minimum residual error but by the elbow criterion, which provides a compromise between information preservation and model complexity.
FIGURE 5
For illustration purposes, the 15-cluster map obtained after defuzzification (Equation 3) is shown in Figure 5b. Clusters are distinguished by color, while color saturation represents the trustworthiness of the crisp cluster assignment (Equation 4). Fully saturated colors indicate map grid nodes with cluster membership values approaching unity.
3.2 Initial route construction
Existing road network data for the survey area are acquired from OpenStreetMap. The filtered road network is shown in Figure 6. The survey area is dominated by track-type roads, complemented by motorways and other important road classes. For CRNS data acquisition, low-speed roads are preferred to ensure robust measurements, as they enable the minimization of spatial integration effects caused by vehicle movement during measurement periods, reduce interference with public traffic, and typically feature simple or unpaved road surfaces.
FIGURE 6
For route construction, we fuse the filtered road network with the clustered auxiliary data. Map grid nodes whose centers are farther than 50 m from any part of the filtered road network are excluded from further analysis and blanked in the clustered map shown in Figure 7. To construct the initial route, information values are assigned as weights to each road segment in the filtered network (Equation 8). For each cluster, the road segment with the lowest weight, corresponding to the highest informational value, is identified. In our application example, this results in the selection of 15 road segments. Pathways connecting all selected segments are computed using Dijkstra’s shortest path algorithm, employing the road segment weights as the cost measure, such that the resulting paths maximize informational value. An ACO algorithm is then used to determine the sequence of visiting the selected segments with minimal total weight. The resulting initial route provides a solution with high informational value, although route length is not yet explicitly optimized.
FIGURE 7
The total length of the initial route is 341 km and is shown in Figure 8. Figure 8a displays the route as a black line over the clustered map previously shown in Figure 5b, illustrating that at least several high-membership grid nodes for each cluster are visited. To quantify normalized inter- and extrapolation distances in the auxiliary data space X, we compute the convex hull of all map grid nodes visited by the initial route. The grid nodes defining the convex hull are used to normalize X column-wise, such that their range spans a unit hypercube (see Figure 2b). Normalized distances are then calculated for each map grid node and are shown in Figure 8b. Approximately 50% of the grid nodes within the survey area lie inside the convex hull defined by the initial route. Normalized interpolation and extrapolation distances are usually well below 0.5 and exceed 1 only for a small number of grid nodes, indicating a relatively good and representative coverage of the survey area by the initial route.
FIGURE 8
3.3 Route economization
Having identified an initial route with high informational value, we now focus on reducing the driving time while preserving its spatial coverage, expressed by low normalized interpolation and extrapolation distances. The convex hull of the initial route in the six-dimensional auxiliary spatial data space comprises 437 vertices (map grid nodes; Figure 9a). These vertices do not necessarily represent all defuzzified clusters (cf.Figure 8a), as clusters located near the center of the data distribution are not required to define the convex hull. With respect to extrapolation, however, the information value provided by these convex hull vertices is essentially equivalent to that of the full route, as extrapolation distances remain nearly unchanged. Interpolation distances increase slightly, because grid nodes on the initial route that lie inside the convex hull are no longer considered when computing interpolation distances (cf.Figures 8b, 9a).
FIGURE 9
Instead of driving slowly along the entire initial route, acquiring CRNS data at all grid nodes, focusing measurements on the 437 convex hull vertices preserves extrapolation fraction and distance while only marginally increasing interpolation distances. This strategy allows driving at regular traffic speeds between these points, improving driving conditions, reducing traffic risks, and lowering overall survey time, while concentrating data acquisition on the most informative locations.
To further reduce the number of sampling points, we identify subsets of points along the initial route that best preserve the volume of the original convex hull. Using PSO, we determine subsets of 100, 50, and 25 nodes that optimally preserve the convex hull volume. Resultant sets of map grid nodes are shown in Figures 9b–d respectively. As the number of retained nodes decreases, both extrapolation fraction and extrapolation distances increase, indicating a progressive loss of information value relative to the full convex hull. We select the solution with 50 nodes as a suitable compromise. This subset preserves approximately 70% of the convex hull volume and keeps extrapolation distances well below 0.5 for nearly all map grid nodes. Using the selected 50 grid nodes, we proceed with route economization. We identify the road segments covering these nodes, compute the shortest-distance connections between them using the A* algorithm, and construct the final route using ACO. The resulting route has a total length of 270 km, compared to 341 km for the initial route. The final route is shown in Figure 10a, together with interpolation and extrapolation distances computed for the full economized route. While this requires driving the entire route slowly to avoid high spatial integration during measurement periods, the route partly follows higher-speed roads between the 50 most informative nodes. As a result, it is shorter, faster to drive, and comparable in information return to the initial route.
FIGURE 10
For comparison and benchmarking in subsequent analyses, we additionally construct an empirical route of similar length (Figure 10b). This route exhibits larger variability in extrapolation distances across the survey area with extensive regions in the southwest and north showing extrapolation distances of approximately 0.5. This indicates a more uneven spatial distribution of information value and a lower overall information return compared to the optimized route.
4 Results and discussion
CRNS data were collected on 30 June 2025, and 1 July 2025, along the optimized and empirically constructed routes, respectively. Both survey days were precipitation-free across the study area. The last rainfall event occurred on 27 June 2025, with amounts ranging from approximately 1 mm in the south to about 4 mm in the northeast. Measurements were conducted using a mobile CRNS system mounted on a vehicle driving along the predefined routes (see for technical details). The primary measurements consist of neutron counts recorded at 12 s intervals along the vehicle trajectory. Following established procedures (; ; ), raw neutron counts were pre-processed using standard corrections for air pressure, air humidity, variations in solar activity, and the presence of roads before conversion to gravimetric soil moisture.
When following the optimized route, minor deviations from the planned trajectory were necessary due to road closures, construction work, and inconsistencies in the OSM database, such as missing or vanished roads. Despite these issues, all 50 targeted map grid nodes were successfully reached. Figure 11a shows the actually driven optimized route together with the 50 most informative points and the corresponding interpolation and extrapolation distances. The total driving time exceeded the planned duration by approximately 3 hours. Contributing factors included missing or inaccurate speed limit information in OSM, poorer-than-expected track conditions, and more than 1 hour spent maneuvering out of dead ends caused by vanished roads, fallen trees, or blocked tracks. Additional delays resulted from frequent turns and traffic-calmed zones in residential areas, which required repeated deceleration of the CRNS vehicle. While using longer road segments and non-urban roads could mitigate this issue in future surveys, such adjustments must be made carefully to avoid disconnected road networks. Human factors also influenced driving time. Trained to drive slowly along entire empirically designed routes in the past, the driver tended not to use the maximum allowable speed between the 50 most informative locations.
FIGURE 11
On the second day, data were collected along the empirically constructed route. Due to operational constraints, this route had to be shortened by approximately 20 km. The actually driven route is shown in Figure 11b, together with normalized interpolation and extrapolation distances computed from measurements along the entire route. In contrast to the optimized route, deviations from the planned trajectory resulted in a data set lacking substantial information coverage in the northern part of the survey area. It should be noted that, in past empirical CRNS survey planning, interpolation and extrapolation distances, or overall information value, were typically not considered, leaving data producers and users unaware of the informational consequences of route modifications. Had the importance of the final section of the route been known, the survey might have been extended by an additional hour or driven at higher speeds.
Figure 12a shows the driving speed along the driven optimized route. Vertical red bars indicate the locations of the 50 most informative road segments, corresponding to the red asterisks in Figure 11a. The driver was instructed to slow down approximately 100 m before entering these segments. This resulted in acquisition speeds below 20 km/h at all critical locations. For further analysis, only CRNS readings collected at speeds of 20 km/h or less were retained. Figure 12b shows the driving speed along the empirically driven route. As no predefined target segments exist for this route, no corresponding markers are shown. Following empirical conventions, measurements acquired at speeds up to 30 km/h were retained for further processing.
FIGURE 12
Figures 12c,d show the CRNS samples acquired at speeds below 20 km/h and 30 km/h along the optimized and empirical routes, respectively, overlaid on normalized distance maps computed using only these measurement locations. Applying these speed thresholds yields 1,184 and 800 accepted CRNS measurements for the optimized and empirical routes, respectively. These data were used to generate spatially continuous soil moisture maps by solving a multiple regression problem with CRNS gravimetric soil moisture measurements as the response variable and the auxiliary maps shown in Figure 4 as predictor variables.
To highlight ontological uncertainty inherent to the regression problem and arising from imperfect informational value of the data sets, we apply two non-parametric regression methods that are equally capable of solving multiple regression problems: random forest (RF) () regression and feed-forward artificial neural networks (ANN) (). Analyzing the differences between the soil moisture maps produced by different regression models for the empirical and optimized routes, we want to mention that RF–ANN soil moisture map disagreement is used as a comparative indicator of model-structure-dependent uncertainty rather than a rigorous uncertainty estimate. The available CRNS measurements provide training data covering 249 and 319 map grid nodes for the optimized and empirical routes, respectively, resulting in a slightly larger training data set for the empirical route. For the optimized route, the training data are built such that all 50 informationally most important grid nodes remain represented in the training data set.
For the optimized route, the gravimetric soil moisture maps derived from the RF and ANN models are shown in Figures 13a,b respectively. Figure 13c shows the difference between the soil moisture maps produced by the two models. The resulting soil moisture maps appear generally realistic, exhibiting consistently higher soil moisture in the Harz Mountains and a narrow zone of increased soil moisture extending from west to east and southeast, corresponding to the Grosses Bruch wetland (cf.Figure 3). Differences between the RF and ANN soil moisture maps are typically below 0.2 and show a weak dependence on soil moisture, with slightly larger differences in the Harz Mountains and along the Grosses Bruch. Localized differences occur mainly in areas characterized by increased extrapolation distances.
FIGURE 13
For the empirical route, the soil moisture maps derived from the RF and ANN regressions and their differences are shown in Figures 13d–f, respectively. Despite the larger training data set, the differences between the two models are more pronounced than for the optimized route. Large areas in the northern and western parts of the survey area exhibit differences exceeding 0.2. These regions coincide with areas of increased extrapolation distances, resulting in increasing differences depending on how the models handle extrapolation, as indicated by the statistical analyses summarized in Table 2. While both models yield generally plausible soil moisture maps given the informational value of the data set, the ontological uncertainty inherent to the regression problem is substantially higher than for the optimized route, and correlated with increased extrapolation requirements (Table 2). While median differences between the soil moisture maps resultant from RF and ANN remain similar for both routes, substantial differences emerge in the upper tail of the disagreement distribution. The 95th percentile increases from 0.23 for the optimized route to 0.38 for the empirical route. Furthermore, RF–ANN disagreement exhibits a stronger correlation with extrapolation distance for the empirical route (ρ = 0.53) than for the optimized route (ρ = 0.35), supporting the hypothesis that increased extrapolation requirements lead to greater model-structure-dependent uncertainty. Consequently, finer spatial soil moisture patterns, such as the elevated moisture zone associated with the Grosses Bruch, are no longer clearly resolved and model-dependent map differences become considerable in parts of the survey area with respect to the actual gravimetric soil moisture.
TABLE 2
| Metric | Optimized route | Empirical route |
|---|---|---|
| Extrapolation fraction | 0.60 | 0.73 |
| Median normalized extrapolation distance | 0.11 | 0.23 |
| Median absolute difference of RF and ANN based soil moisture map | 0.06 | 0.07 |
| 95th percentile of absolute difference of RF and ANN based soil moisture map | 0.23 | 0.38 |
| Spearman correlation between normalized extrapolation distance and the difference between RF and ANN soil moisture maps | 0.35 | 0.53 |
Quantitative comparison of the optimized and empirical survey routes. Extrapolation metrics characterize the representativeness of the acquired data with respect to the auxiliary-data space.
Differences between Random Forest (RF) and Artificial Neural Network (ANN) soil moisture maps are used as a comparative indicator of model-structure-dependent (ontological) uncertainty. Spearman correlation coefficients quantify the relationship between normalized extrapolation distance and RF–ANN disagreement.
Future work should investigate the performance of the proposed framework relative to alternative survey-design strategies, including random, uniform, and information-theoretic sampling schemes. While the present study focuses on demonstrating the practical benefits of information-return-driven survey design compared to empirically planned field campaigns, a systematic benchmarking study could provide additional insight into the relative strengths and limitations of different route construction and sampling concepts under varying environmental and logistical conditions.
5 Conclusion
We have developed a general solution framework for maximizing the informational return of mobile data acquisition systems operating on road networks. The approach fuses multiple spatially dense auxiliary data sets that describe the spatial heterogeneity of a survey area with road network information and formulates route design as a combinatorial optimization problem. To ensure computational tractability while preserving information content, the framework follows a two-step strategy. First, an initial route is constructed that is guaranteed to be information rich by explicitly accounting for spatial heterogeneity. In a second step, this route is economized by identifying a reduced set of critical sampling locations that preserves the information content of the initial solution, as quantified through convex hull properties in the auxiliary data space. Direct application of the economization step to the full road network would be computationally infeasible, which makes the proposed two-step approach essential.
The effectiveness of the framework was demonstrated using a large-scale case study in central Germany, where soil moisture was measured using a vehicle-mounted Cosmic Ray Neutron Sensing (CRNS) system. Field measurements along both optimized and empirically designed routes confirm the theoretical findings. The optimized route achieved a substantially higher and more spatially balanced informational return. Soil moisture maps derived from the optimized route exhibited lower extrapolation requirements and reduced ontological uncertainty in subsequent regression-based regionalization compared to maps obtained from an empirically constructed route comprising a larger training data set when learning the regression models.
Beyond the specific application presented here, the proposed methodology is broadly applicable to mobile sensing problems in Earth and environmental sciences where sparse measurements must be collected efficiently over large areas. By explicitly linking survey design, information theory, and optimization, the framework provides a systematic alternative to empirically driven or convenience-based sampling strategies and offers a transparent means to quantify and compare the informational value of alternative survey designs.
Statements
Data availability statement
The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.
Author contributions
CT: Conceptualization, Formal Analysis, Funding acquisition, Methodology, Software, Visualization, Writing – original draft. LT: Data curation, Software, Writing – review and editing. JA: Data curation, Software, Supervision, Writing – review and editing. SL: Investigation, Writing – review and editing. MS: Data curation, Funding acquisition, Resources, Writing – review and editing. PD: Conceptualization, Supervision, Writing – review and editing. HP: Conceptualization, Investigation, Project administration, Resources, Validation, Writing – review and editing.
Funding
The author(s) declared that financial support was received for this work and/or its publication. CT was supported by the Study Abroad Program of The Republic of Türkiye Ministry of National Education. The research was supported by the Deutsche Forschungsgemeinschaft (Grant 357874777; research unit FOR 2694, Cosmic Sense II).
Conflict of interest
The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Generative AI statement
The author(s) declared that generative AI was not used in the creation of this manuscript.
Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
References
1
AltdorffD.OswaldS. E.ZachariasS.ZengerleC.DietrichP.MollenhauerH.et al (2023). Toward large-scale soil moisture monitoring using rail-based cosmic ray neutron sensing. Water Resour. Res.59 (3), e2022WR033514. 10.1029/2022WR033514
2
AndradeC. (2020). Sample size and its importance in research. Indian J. Psychol. Med.42 (1), 102–103. 10.4103/IJPSYM.IJPSYM_504_19
3
AroraJ. S. (2012). Introduction to Optimum Design. Academic Press. 10.1016/C2013-0-15344-5
4
BabaeianE.SadeghiM.JonesS. B.MontzkaC.VereeckenH.TullerM. (2019). Ground, proximal, and satellite remote sensing of soil moisture. Rev. Geophys.57 (2), 530–616. 10.1029/2018RG000618
5
BaranT.HarmanciogluN. B.CetinkayaC. P.BarbarosF. (2017). An extension to the revised approach in the assessment of informational entropy. Entropy19 (12), 634. 10.3390/e19120634
6
BaroniG.ScheiffeleL. M.SchrönM.IngwersenJ.OswaldS. E. (2018). Uncertainty, sensitivity and improvements in soil moisture estimation with cosmic-ray neutron sensing. J. Hydrology564, 873–887. 10.1016/j.jhydrol.2018.07.053
7
BianchessiN.CorberánÁ.PlanaI.ReulaM.SanchisJ. M. (2022). The profitable close-enough arc routing problem. Comput. Operations Res.140, 105653. 10.1016/j.cor.2021.105653
8
BreimanL. (2001). Random forests. Mach. Learn.45, 5–32. 10.1023/A:1010933404324
9
BrownW. G.CoshM. H.DongJ.OchsnerT. E. (2023). Upscaling soil moisture from point scale to field scale: toward a general model. Vadose Zone J.22 (2), e20244. 10.1002/vzj2.20244
10
ChrismanB.ZredaM. (2013). Quantifying mesoscale soil moisture with the cosmic-ray rover. Hydrology Earth Syst. Sci.17, 5097–5108. 10.5194/hess-17-5097-2013
11
CuiH.SivakumarB.SinghV. P. (2018). Entropy application in environmental and water engineering. Entropy20, 598. 10.3390/e20080598
12
DegaS.DietrichP.SchrönM.PaascheH. (2023). Probabilistic prediction by means of the propagation of response variable uncertainty through a monte carlo approach in regression random forest: application to soil moisture regionalization. Front. Environ. Sci.11, 1009191. 10.3389/fenvs.2023.1009191
13
DijkstraE. W. (1959). A note on two problems in connexion with graphs. Numer. Math.1, 269–271. 10.1007/BF01386390
14
DobbieM. J.HendersonB. L.Stevens, JrD. L. (2008). Sparse sampling: spatial design for monitoring stream networks. Stat. Surv.2, 113–153. 10.1214/07-SS032
15
FanK. (1959). Convex Sets and their Applications.
16
FerschB.JagdhuberT.SchrönM.VölkschI.JägerM. (2018). Synergies for soil moisture retrieval across scales from airborne polarimetric SAR, cosmic ray neutron roving, and an in situ sensor network. Water Resour. Res.54 (11), 9364–9383. 10.1029/2018WR023337
17
FineT. L. (1999). Feedforward Neural Network Methodology. Springer.
18
FranzT. E.WangT.AveryW.FinkenbinerC.BroccaL. (2015). Combined analysis of soil moisture measurements from roving and fixed cosmic ray neutron probes for multiscale real-time monitoring. Geophys. Res. Lett.42 (9), 3389–3396. 10.1002/2015GL063963
19
FranzT.LariosA.VictorC. (2022). The bleeps, the sweeps, and the creeps: convergence rates for dynamic observer patterns via data assimilation for the 2D navier–stokes equations. Comput. Methods Applied Mechanics Engineering392, 114673. 10.1016/j.cma.2022.114673
20
FuchsF.KolínskýP.GröschlG.ApolonerM. T.QorbaniE.SchneiderF.et al (2015). Site selection for a countrywide temporary network in Austria: noise analysis and preliminary performance. Adv. Geosciences41, 25–33. 10.5194/adgeo-41-25-2015
21
GambardellaL. M.DorigoM. (1996). Solving symmetric and asymmetric TSPs by ant colonies. IEEE Conf. Evol. Comput., 622–627. 10.1109/ICEC.1996.542672
22
GoldenB. L.LevyL.VohraR. (1987). The orienteering problem. Nav. Res. Logist.34 (3), 307–318. 10.1002/1520-6750(198706)34:3<307::AID-NAV3220340302>3.0.CO;2-D
23
HathawayR. J.BezdekJ. C. (2001). Fuzzy c-Means clustering of incomplete data-part B: cybernetics. IEEE Trans. Syst. Man Cybern.31 (5), 735–744. 10.1109/3477.956035
24
HawdonA.McJannetD.WallaceJ. (2014). Calibration and correction procedures for cosmic-ray neutron soil moisture probes located across Australia. Water Resour. Res.50 (6), 5029–5043. 10.1002/2013WR015138
25
JakobiJ.HuismanJ. A.VereeckenH.DiekkrügerB.BogenaH. R. (2018). Cosmic ray neutron sensing for simultaneous soil water content and biomass quantification in drought conditions. Water Resour. Res.54, 7383–7402. 10.1029/2018WR022692
26
JiangL.TianG.WangB.GuoX.HeX.ZouA.et al (2021). Application of three-dimensional electrical resistivity tomography in urban zones by arbitrary electrode distribution survey design. J. Appl. Geophys.194, 104460. 10.1016/j.jappgeo.2021.104460
27
KennedyJ.EberhartR. (1995). “Particle swarm optimization,” (Perth: IEEE). Proceedings on ICNN'95 - International Conference on Neural Networks10.1109/ICNN.1995.488968
28
KiureghianA. D.DitlevsenO. (2009). Aleatory or epistemic? Does it matter?Struct. Saf.31, 105–112. 10.1016/j.strusafe.2008.06.020
29
KöhliM.SchrönM.ZredaM.SchmidtU.DietrichP.ZachariasS. (2015). Footprint characteristics revised for field-scale soil moisture monitoring with cosmic-ray neutrons. Water Resour. Res.51, 5772–5790. 10.1002/2015WR017169
30
LiH.WangD.SinghV. P.WangY.WuJ.WuJ. (2021). Developing an entropy and copula-based approach for precipitation monitoring network expansion. J. Hydrology598, 126366. 10.1016/j.jhydrol.2021.126366
31
LowK. H.DolanJ. M.KhoslaP. (2013). “Information-theoretic approach to efficient adaptive path planning for Mobile robotic environmental sensing,”19th International Conference on Automated Planning and Scheduling (ICAPS 2009). 10.48550/arXiv.1305.6129
32
MaurerH.BoernerD. E.CurtisA. (2000). Design strategies for electromagnetic geophysical surveys. Inverse Probl.16 (5), 1097–1117. 10.1088/0266-5611/16/5/302
33
MuX.HuM.SongW.RuanG.GeY.WangJ.et al (2015). Evalutaion of sampling methods for validation of remotely sensed fractional vegetation cover. Remote Sens.7, 16164–16182. 10.3390/rs71215817
34
NoshadM.ChoiJ.SunY.HeroA.DinovI. D. (2021). A data value metric for quantifying information content and utility. J. Big Data8 (82), 82. 10.1186/s40537-021-00446-6
35
OleaR. A. (1984). Sampling design optimization for spatial functions. J. Int. Assoc. Math. Geol.16, 369–392. 10.1007/BF01029887
36
PaascheH.EberleD. (2011). Automated compilation of pseudo-lithology maps from geophysical data sets: a comparison of gustafson-kessel and fuzzy c-means cluster algorithms. Explor. Geophys.42 (4), 275–285. 10.1071/EG11014
37
PaascheH.DegaS.SchrönM.DietrichP. (2025). Comprehensive data aleatory uncertainty propagation in regression random forest using a monte carlo apprach: a struggle with incomplete data provision using a case study on probabilistic soil moisture regionalization. Front. Environ. Sci.13, 1599320. 10.3389/fenvs.2025.1599320
38
PopovićM.Vidal-CallejaT.HitzG.ChungJ. J.SaI.SiegwartR.et al (2020). An informative path planning framework for UAV-Based terrain monitoring. Aut. Robots44, 889–911. 10.1007/s10514-020-09903-2
39
RussellS. J.NorvigP. (2020). Artificial Intelligence. Pearson Education Limited.
40
SchmidtK.BehrensT.DaumannJ.Ramirez-LopezL.WerbanU.DietrichP.et al (2014). A comparison of calibration sampling schemes at the field scale. Geoderma232-234, 243–256. 10.1016/j.geoderma.2014.05.013
41
SchrönM.RosolemR.KöhliM.PiussiL.SchröterI.IwemaJ.et al (2018). Cosmic-ray neutron rover surveys of field soil moisture and the influence of roads. Water Resour. Res.54 (9), 6441–6459. 10.1029/2017WR021719
42
SchrönM.OswaldS. E.ZachariasS.KasnerM.DietrichP.AttingerS. (2021). Neutrons on rails: transregional monitoring of soil mooisture and snow water equivalent. Geophys. Res. Lett.48 (24), e2021GL093924. 10.1029/2021GL093924
43
SchröterI.PaascheH.DietrichP.WollschlägerU. (2015). Estimation of catchment-scale soil moisture patterns based on terrain data and sparse TDR measurements using a fuzzy C-Means clustering approach. Vadose Zone J.14 (11), 1–16. 10.2136/vzj2015.01.0008
44
SinghV. P. (2013). Entropy-Based Parameter Estimation in Hydrology. Dordrecht: Springer.
45
StummerP.MaurerH.GreenA. G. (2004). Experimental design: electrical resistivity data sets that provide optimum subsurface information. Geophysics69 (1), 120–139. 10.1190/1.1649381
46
TaylorJ. R. (1982). An Introduction to Error Analysis. Sausalito, CA:University Science Books.
47
Van LeekwijckW.KerreE. E. (1999). Defuzzification: criteria and classification. Fuzyy Sets Syst.108 (2), 159–178. 10.1016/S0165-0114(97)00337-0
48
van WestenC. J.CastellanosE.KuriakoseS. L. (2008). Spatial data for lanslide susceptibility, hazard and vulnerability assessment: an overview. Eng. Geol.102, 112–131. 10.1016/j.enggeo.2008.03.010
49
WangZ.ZhengJ.HanC.LuB.YuD.YangJ.et al (2024). Exploring the potential of OpenStreetMap data in regional economic development evaluation modeling. Remote Sens.16 (2), 239. 10.3390/rs16020239
50
WilkinsonP. B.LokeM. H.MeldrumP. I.ChambersJ. E.KurasO.GunnD. A.et al (2012). Practical aspects of applied optimized survey design for electrical resistivity tomography. Geophys. J. Int.189, 428–440. 10.1111/j.1365-246X.2012.05372.x
51
WilliamsK. J.BelbinL.AustinM. P.SteinJ. L.FerrierS. (2012). Which environmental variables should I use in my biodiversity model?Int. J. Geogr. Inf. Sci.26 (11), 2009–2047. 10.1080/13658816.2012.698015
52
WollschlägerU.HelmingK.HeinrichU.BartkeS.Kögel-KnabnerI.RussellD.et al (2016). “The BonaRes Centre-A virtual institute for soil research in the context of a sustainable bio-economy,” in Geophysical Research Abstracts, 18. Vienna: European Geosciences Union.
53
YangY.YangJ. J. (2023). Strategic sensor placement in expansive highway networks: a novel framework for maximizing information gain. Systems, 11 (12), 577. 10.3390/systems11120577
54
YangY.ZhenH.YangJ. J. (2024). An information gradient approach to optimizing traffic sensor placement in statewide networks. Information, 15(10), 654. 10.3390/info15100654
55
ZachariasS.LoescherH. W.BogenaH.KieseR.SchrönM.AttingerS.et al (2024). Fifteen years of integrated terrestrial environmental observatories (TERENO) in Germany: functions, services, and lessons learned. Earth's Future12 (6), e2024EF004510. 10.1029/2024EF004510
56
ŽížalaD.PrincT.SkálaJ.JuřicováA.LukasV.BohovicR.et al (2024). Soil sampling design matters - enhancing the efficiency of digital soil mapping at the field scale. Geoderma Reg.39, e00874. 10.1016/j.geodrs.2024.e00874
Summary
Keywords
cluster analysis, combinatorial optimization, cosmic ray neutron sensing, experimental design, information fusion, soil moisture, vehicle routing
Citation
Topaclioglu C, Trinkle L, Anders J, Landmark S, Schrön M, Dietrich P and Paasche H (2026) Maximizing the information return of vehicle-based sensors: optimal routing for mobile cosmic ray neutron sensing. Front. Environ. Sci. 14:1803901. doi: 10.3389/fenvs.2026.1803901
Received
04 February 2026
Revised
11 June 2026
Accepted
15 June 2026
Published
07 July 2026
Volume
14 - 2026
Edited by
Juergen Pilz, University of Klagenfurt, Austria
Reviewed by
Michael Tso, UK Centre for Ecology and Hydrology (UKCEH), United Kingdom
Yunxiang Yang, University of Georgia, United States
Updates
Copyright
© 2026 Topaclioglu, Trinkle, Anders, Landmark, Schrön, Dietrich and Paasche.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Can Topaclioglu, can.topaclioglu@ufz.de
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.