ORIGINAL RESEARCH article

Front. Photonics, 18 February 2026

Sec. Neuromorphic Photonics and Photonic Computing

Volume 7 - 2026 | https://doi.org/10.3389/fphot.2026.1696425

Wavelength selection for laser design in mid-infrared spectroscopy

  • 1. Faculty of Science and Technology, Norwegian University of Life Sciences, Ås, Norway

  • 2. Department of Chemistry, University of Rome La Sapienza, Rome, Italy

  • 3. Nofima AS - Norwegian Institute of Food, Fisheries and Aquaculture Research, Ås, Norway

  • 4. Department of Physical and Chemical Sciences, University of L’Aquila, L’Aquila, Italy

Abstract

The development of miniaturized tunable laser sources for mid-infrared (MIR) spectroscopy has enabled portable, application-specific analytical devices. Recent advances in quantum cascade lasers (QCLs) and interband cascade lasers (ICLs) allow precise wavelength emission in narrow spectral regions, such as 1700–1600 , which is critical for protein characterization. In this study, we evaluate machine learning techniques for selecting the most informative wavelengths to guide the design of tunable laser systems, and for their ability to account for specific constraints such as the possibility to do fine and coarse laser wavelength tuning. We focus on optimizing variable selection for a laser-based device targeting peptide analysis and protein quality assessment in hydrolysates as a case study. We compare sparse modelling techniques (SPLS), filter-based (SPA, CovSel, g-CovSel), and compression methods (PVS, PVR), and propose a new algorithm (w-CovSel) to assess their ability to reduce noise and isolate key spectral features. Our results highlight the potential of providing data-driven approaches to obtain laser design which enables high-performance MIR instrumentation tailored to specific analytical tasks.

1 Introduction

Lasers have revolutionized spectroscopy, becoming powerful radiation sources thanks to advances in solid-state physics (). Their high intensity enhances the signal-to-noise ratio (SNR) and sensitivity, thus enabling measurements of higher excited states. Moreover, it leads to shorter acquisition times, facilitating faster imaging measurements and more precise studies of dynamic processes. Finally, their exceptional spectral resolution offers detailed structural insights (). These characteristics have made lasers integral to various spectroscopic techniques such as absorption, fluorescence, photothermal, photoacoustic, and Raman, across spectral regions from the ultraviolet to the far-infrared (; ). Visible (e.g., red, green) and near-infrared (NIR) lasers are commonly used as light sources in Raman spectroscopy (), whereas semiconductor diode lasers are widely employed for gas monitoring and detection (; ; ).

Tunable lasers offer superior SNR, sensitivity, and selectivity (), and recent developments have positioned them as promising alternatives to traditional incoherent sources in Fourier-Transform Infrared (FTIR) spectrometers, facilitating miniaturization (; ; ). Among these, quantum cascade lasers (QCLs) and interband cascade lasers (ICLs) have emerged as key technologies in mid-infrared (MIR) spectroscopic sensing over the past decade (; ; ; ). Commercial QCLs (e.g., MIRCAT, LaserTune) are widely used in MIR spectrometers and microscopes (; ), while ICLs are particularly suited for portable applications due to their efficient operation within narrow MIR ranges (). Both laser types can be tailored to emit specific wavelength ranges, supporting highly versatile and compact systems. ICLs, in particular, enable continuous-wave (CW) operation at room temperature across the 3.0–6.5 wavelength range, with low power consumption, single-mode operation, and precise modulation capabilities (; ; ; ). Coarse tuning can span several hundred wavenumbers (wns), while fine tuning achieves sub-wavenumber precision, typically well within 1 , by adjusting temperature and current (; ).

Recent works have demonstrated the use of single-wavelength and tunable lasers in the 5.88–6.09 (1700–1640 ) range for targeted applications in food safety, where such a laser has been used in a device for mycotoxins detection (; ). This same spectral region is highly relevant for other biological application, such as protein analysis, since the amide I band is positioned in this range, thus significantly broadening the scope of these lasers’ applications. The amide I band, corresponding to C=O stretching in the peptide backbone, is known to be sensitive to protein secondary structure and conformational changes (). The state-of-the-art tunable laser mentioned above, can ideally emit at over 1000 discrete frequencies (in theory there is not a limit), however, the optimized ones chosen for this application are: 1691.98, 1691.49, 1690.23, 1689.04, 1688.42, 1673.79, 1673.3, 1672.45, 1667.61, 1667.37, 1665.93, 1665.41, 1664.49, 1651.78, 1651.48, 1650.5, 1649, 1642.05, 1641.47, and 1640.49 . As shown in the list, some wavelengths offer narrow tunable windows, while others are more sparsely distributed. This enables the construction of simple, compact measurement setups requiring only simplistic beam guiding optics (tunable laser, sampling unit, and detector), bypassing the optical complexity of traditional FTIR systems.

Laser design critically depends on selecting the most informative wavelengths. Importantly, tunability is subject to physical and technical constraints: material properties and manufacturing limitations can restrict both the coarse and fine tuning ranges (; ; ). Therefore, selecting a small and well-distributed set of critical wavelengths, a process known as variable selection in machine learning, is a central challenge in the design of such lasers.

In spectroscopy, variable selection is essential due to the high dimensionality of spectral data, which often contains redundant or noisy variables. The primary goal is to enhance interpretability and improve performance by isolating the most relevant features. Variable selection methods can be broadly categorized into filter, wrapper, and embedded approaches, depending on how variable selection interacts with model building. Filter methods operate independently of the model and allow stricter control over the number of selected variables, which is an important feature in laser design. Wrapper methods balance model performance and variable choice, while embedded methods integrate selection into model training (). Classical approaches, like Principal Components Analysis (PCA) and Partial Least Squares Regression (PLSR) remain widely used, but more advanced techniques such as Least Absolute Shrinkage and Selection Operator (LASSO; ()), Uninformative Variables Elimination (UVE; ()), and Variable Importance Projection (VIP; ()) are also widely applied. Recent developments include more complex machine learning methods, such as neural networks, genetic algorithms, and random forests (; ; ; ; ).

The aim of this study was to evaluate different variable selection methods, with particular emphasis on their applicability to the design of tunable infrared lasers that require both coarse and narrow tuning. A miniaturized laser tunable within the 1700–1600 spectral region (; ) exemplifies such a system, offering both wide and fine tuning capabilities. Accordingly, variable selection techniques must be able to identify not only discrete informative wavelengths, but also narrow spectral regions. Thus, the main objective was to identify methods that can robustly select a compact set of informative wavelengths while considering the practical constraints of laser tunability, such as wide and fine tuning.

For the study, we primarily focused on filter-based methods like the Successive Projection Algorithm (SPA), and Covariance Selection (CovSel), along with its extensions for groups (g-CovSel) and windows (w-CovSel, newly implemented) of variables. An unsupervised method called Principal Variable Selection (PVS) was included along with it supervised extension Principal Variable Regression (PVR). While embedded and wrapper methods often select too many variables and struggle with collinearity, one notable exception is Sparse Partial Least Squares (SPLS), which allows explicit control of sparsity (e.g., 90% variable reduction) (; ; ). Thus, this method was also taken in consideration.

2 Materials and methods

2.1 Data

This study utilizes FTIR spectra of protein hydrolysate samples derived from previous research conducted at Nofima AS (Ås, Norway) (; ). Ethical approval was not required for these studies, as the samples originated from salmon and poultry by-products obtained through routine food industrial processes.

To explore optimal laser emission wavelengths, we analyzed broadband infrared data acquired through benchtop Fourier-Transform Infrared (FTIR) instrumentation. The variables in the datasets span the full mid-infrared (mid-IR) range from 4000 to 400 , corresponding to wavelengths of 2.5–25 , with a spectral resolution of 4 and a 5.0 mm aperture, as described by ().

This study focused on three distinct datasets, all derived from FTIR analysis of protein ingredients obtained via enzymatic protein hydrolysis, a process involving enzymatic breakdown of protein-rich raw materials into smaller protein and peptide components (i.e., protein hydrolysates). These products are enriched in peptides and free amino acids, with their average molecular weight (AMW) crucial for determining product quality, functional properties, bioavailability, and sensory characteristics. The AMW of peptides varies based on sample type, enzymes used, and process conditions.

The first dataset comprises 95 samples measured in one technical replicate of salmon hydrolysates (Fish dataset), prepared in a laboratory setting through enzymatic hydrolysis as described by (). Additionally, two datasets of poultry hydrolysates were collected at the Bioco facility in Hærland, Norway. These include 133 samples measured in five technical replicates (), and 286 samples measured in three technical replicates (), obtained under varying processing conditions over different time periods at the facility. These datasets are referred to as Poultry 1 (133 samples) and Poultry 2 (286 samples).

The AMW average values and standard deviation for the datasets are as follows: 1885 214 g/mol for Fish dataset, 4858 910 g/mol for Poultry 1 dataset, and 7111 2176 g/mol for Poultry 2 dataset. In the Poultry 2 dataset, the higher standard deviation can be attributed to the experimental design, which involved extensive variations in raw materials, enzyme types, and water flow rates in the enzymatic process lines. These values were determined via size-exclusion chromatographic analysis of the respective samples, consistent with previous reports (; ).

2.2 Variable selection

Proteins and peptides exhibit characteristic infrared absorption bands (amide A, B, and I-VII). Among these, the fingerprint range (1800–800 ) is particularly rich in information, capturing vibrational modes that reveal details on protein backbones, secondary structures, and chemical composition. Based on literature, a set of wavenumbers (wns) 1740 (lipids), 1645, 1630, 1548, 1515, 1452, 1400, 1081, and 1045 (hereafter referred to as Set of biochemically motivated features) has been identified as particularly relevant for characterizing biomolecular composition, for instance proteins and peptides (; ; ). The 1740 signal remains unaffected by protein absorptions and therefore, was used as a reference to account for baseline variations. FTIR analysis of protein hydrolysates obtained from enzymatic hydrolysis of rest raw materials from poultry processing showed that amide I (1650 ) and amide II (1550 ) are the most prominent. Additional infrared bands become evident, including absorption near 1516 ( deformation at amino terminals) and 1400 ( groups at carboxyl terminals) (; ), while the region around 1450 corresponds to the amide III band. Further absorption bands at 1045 and 1081 are linked to skeletal vibrations, collagen structures, and C-O stretching ().

For the purpose of laser development, the wns included in the set of biochemically motivated features were evaluated alongside those already accessible in the current laser system: 1691.38, 1689.46, 1687.53, 1674.03, 1672.10, 1668.24, 1666.31, 1664.38, 1650.88, 1648.95, and 1641.24 (selected from larger wns dataset considering the resolution of the available dataset from FTIR spectra). The key hardware specifications of the existing laser are listed in Supplementary Material. In addition, a range of variable selection methods were tested, which are described in the following sections.

When considering very sparse data with wavelengths spread across a broad spectral region, it would be unrealistic to rely on a high number of single-emission points, i.e., single-wavelength lasers. Recently, an instrument designed for in vivo cartilage analysis, containing seven distinct single-emission lasers, has been built and validated (; ). Therefore, the maximum number of variables to select was set to ten, when possible, which is also a sufficient number of variables for good regression modelling on mid-infrared data ().

Prior to the selection through the methods described below, technical replicates were averaged when present and the datasets were split: the Fish and Poultry 1 datasets were randomly divided into a 70% calibration set used for variable selection, eventually used as a training set to predict the AMW of the remaining 30% of the samples (independent test set). For the Poultry 2 dataset, which includes samples processed using different enzymes, the training: test split was conducted using a stratified sampling strategy with three folds. This approach ensured an equal representation of samples from each enzyme across all subsets. The first fold was retained as an independent test set for AMW regression, while the remaining two folds were used both for variable selection and as the calibration set for regression.

Classical Least Squares (CLS) could not be applied in this study due to the inability to isolate pure absorbance spectra for the sample constituents (; ). Instead, Partial Least Squares Regression (PLSR) was utilized, as it has been the standard approach for spectroscopic data analysis of complex matrices since the 1980s (; ). The number of components was set to 10. The results are reported in terms of coefficient of determination (R2), and root mean squared error (RMSE) both calculated in prediction on the independent test set. The data were only pre-processed by mean centering before the variable selection and regression steps. The chosen variable selection methods differ in the criteria and mechanisms used to identify the variables to select. Some techniques select isolated variables, while others focus on groups or narrow ranges of variables. In Figure 1 a schematic representation of the design workflow is reported.

FIGURE 1

2.2.1 Sparse Partial Least Squares (SPLS)

SPLS is a technique that performs variable selection and dimensionality reduction simultaneously, as an integrated part of the modelling process, and belongs to the category of so-called embedded methods. In these approaches, variable selection is carried out at the component level, where variables with smaller loading weights are penalized. The goal is to find a sparse weight vector w, such that the latent score t = Xw has maximal covariance with y. The problem can then be formulated as shown in Equation 1:

subject to:

with being the regularization coefficient enforcing sparsity, by shrinking many coefficients in w to zero, as shown in Equation 2. The algorithm typically initializes w as y and then applies soft-thresholding to obtain a sparse w, normalized to satisfy the L2 constraint. X and y are then deflated to extract further components if needed. This formulation ensures that the method effectively handles multicollinearity, making it particularly suitable for spectroscopic data, as demonstrated by (; ; ).

In this study, the method was applied with a total of 3 components and varying sparsity degrees tailored to different spectral ranges. For full (4000–400 , a total of 1866 initial wns), and for fingerprint (1800–800 , 519 wns) ranges a sparsity degree of 0.99 was used; for the amide I range (1700–1600 , with a total of 52 wns), a lower sparsity degree of 0.90 was chosen. This adjustment was necessary due to (i) the large amount of variables in the full and fingerprint ranges, and at the same time (ii) the limited number of them in the amide I range, where a higher sparsity degree would result in an impractically small number of selected variables. This choice of sparsity degrees would ideally result in 19, 6, and 8 selected variables for full, fingerprint, and amide I ranges, respectively. However, these numbers are often not the ones that we obtain due to the fact that the method does not shrink the coefficients to zero exactly, but could carry some small information, removing, e.g., less than 99% of the variables initially present in the full range.

2.2.2 Principal Variable Selection and regression (PVS and PVR)

PVS is an unsupervised variable selection method that neither filters, wraps nor embeds into another method. The core principle of PVS is a voting strategy that identifies the variable in X, explaining most of the variation in the dataset before deflating and repeating, i.e., a single variable version of Principal Component Analysis (PCA). This iterative orthogonal process allows for the identification and selection of the variables that contribute most significantly to the overall variance of the dataset. If the variables are highly correlated, as in many spectroscopic techniques, the results are similar to PCA components (PCs) (; ; ; ).

The PVS method can be extended to a supervised regression context, known as Principal Variable Regression (PVR). By introducing a response matrix, Y, the variable selection process is done on a joint criterion for the explanation of X and Y. Here, the focus shifts to selecting variables that exhibit the maximum correlation between the predictor matrix (X) and the response matrix (Y). This approach ensures that the selected variables not only explain variance but also have strong predictive relationships with the response variable (). The PVS and PVR processes described here were implemented with 3 PCs, and 10 as maximum number of variables to select. Ideally, an optimization by cross-validation is preferred to choose the number of variables, but this number was chosen as explained in the previous Variable Selection Section 2.2.

2.2.3 Successive Projection Algorithm (SPA)

SPA is a forward variable selection technique which implements an iterative process. The algorithm is designed to minimize collinearity among the selected variables, making it particularly effective in handling datasets with strong multicollinearity. The number of variables to be selected is set manually, allowing for flexibility based on the needs of the analysis. The selection process begins by identifying the first variable with the largest norm (magnitude) in the predictor matrix (X). This ensures that the first selected variable corresponds to the highest variance in the dataset. Mathematically, this step is expressed by Equation 3:

where represents the ith column of X. Subsequent variables are selected based on maximizing orthogonality through the projection norm. Specifically, each candidate variable is adjusted to account for the collinearity with the previously selected variables through the following Equation 4:

where represents the matrix of variables already selected. This iterative process ensures that the newly selected variables are as orthogonal as possible to the previously chosen ones. The method thus prioritizes variables that contribute unique and complementary information to the model, as described in studies by (; ; ).

In this study, the technique was applied without autoscaling, meaning the variables were not normalized prior to selection. This approach preserves the original scale of the data, which may be advantageous in certain applications where variable magnitudes carry meaningful information.

2.2.4 Covariance Selection (CovSel) family

2.2.4.1 CovSel

CovSel is a filter-based selection method that identifies the variables which have the highest covariance with the response variable (Y). The primary objective is to select a subset of variables that maximize the covariance with Y, ensuring that the chosen variables contribute the most to explaining the dependent variable. The algorithm begins by calculating the squared covariance between each predictor variable (x) and the response variable (y). Assuming that both x and y are mean centered, this is expressed mathematically as expressed by Equation 5:

The variables are ranked based on their covariance values, and the highest-ranking variables are iteratively selected. The procedure is iterated after orthogonalizing both X and Y with respect to the selected variables. This process not only ensures that the chosen variables are strongly correlated with the response variable but also helps reduce redundancy among the selected variables by focusing on those that provide unique contributions. The final number of desired variables must be specified by the user prior to running the algorithm. Before the analysis, the data are centered to remove the influence of offsets.

In this study, the maximum number of variables to select was set to 10. The algorithm bears similarities to the supervised component of PVS (PVR, see Methods Section 2.2.2). However, a key distinction lies in the selection criteria. In this method, the selected variables tend to be both highly correlated with the response variable and possess a high norm (magnitude). In contrast, in PVR, the selection prioritizes variables that are highly correlated with the response variable but also align with PCs that have larger singular values. This subtle difference affects the selection process, with the current method favoring variables that exhibit both high correlation and strong individual significance in the dataset (; ).

2.2.4.2 Group(g)-CovSel

G-CovSel is a selection technique based on similar principles to those of CovSel, but is designed to select one or more groups of variables rather than individual variables (). The selection process begins with the identification of a “seed” variable, which serves as the starting point for building the variable groups. This variable is chosen based on the same principle as the original CovSel, i.e., the one having the maximum covariance with the response. Then, the selection of a group of variables with similar characteristics relies on identifying predictors having simultaneously a high correlation with the seed and a high covariance with the response: this is accomplished by setting threshold values for both the correlation coefficient with the seed () and normalized absolute value of covariance with the response (; where normalization is carried out with respect to the covariance of the seed, to have values scaled between 0 and 1), so that only predictors with both values above the respective thresholds are included in the group. Accordingly, if the initial group of selected variables is too large, stricter threshold values can be applied to reduce the number of selected variables. Conversely, if the selected group size is too small, the thresholds can be relaxed to include more variables in the group. The process is then iterated after deflation of both X and Y with respect to the selected variables until a predefined number of groups is selected.

This flexibility in threshold adjustment allows the method to adapt to different datasets and objectives, making it particularly useful in cases where the relationships between variables and the response variable (Y) are complex or when there is a need to balance between selecting a compact set of variables and capturing sufficient information. By focusing on covariance within the data, this technique provides a straightforward yet powerful approach for identifying groups of variables that are collectively relevant to the response variable (). The present work adopted five groups to select, and threshold values of = 0.95 and = 0.7 (the default values are and = 0.8). These values have been chosen to obtain a reasonable amount of selected laser wavelengths.

2.2.4.3 Window(w)-CovSel

W-CovSel is a variable selection technique which we specifically designed to deal with laser fine tuning, as it focuses on the possibility of selecting windows of adjacent predictors, e.g., for spectroscopic applications. As the name suggests, w-CovSel extends the principles of CovSel, to the selection of fixed width windows of variables. After having decided the number of consecutive variables constituting the interval (the window size), the first interval is selected as the one having the highest squared covariance with the response. Then, both X and Y are deflated with respect to the selected interval (and the deflation can be either full rank or rank-1, the latter corresponding to the first singular vector of the variables in the interval), and the procedure is iterated until a predefined number of intervals is selected. The optimal number of intervals is usually determined by cross-validation.

In the present study, the maximum number of intervals was set to five for the full and fingerprint ranges, and to three for the 1700–1600 (amide I) range. Fine tuning in lasers is easy to achieve (), and we have previously shown that the fine tuning of selected laser emission points gives considerably better prediction ability than only using the same emission points without that. Optimal performances were already achieved starting from only 3 emission wavelengths with 3 tuned points each (), and therefore the window size was set to 3 consecutive variables within the selected intervals.

It is important to note that fine tuning can be done for single wavelength lasers. The new laser can be coarsely tuned in the wavelength range 1700–1600 (). Therefore, it makes sense to consider fine tuning around single wavelength emission points in the full mid-IR range, in the fingerprint range and in the 1700–1600 range.

3 Results

3.1 Using variables from key spectral regions and existing tunable laser

The average FTIR spectra of three datasets used in this paper are shown in Figure 2, i.e., the Fish hydrolysates (Figure 2a), Poultry 1 (Figure 2b), and Poultry 2 (Figure 2c) hydrolysates datasets. The Fish dataset display distinct amide band features, mainly in fingerprint (1800–800 ) range, compared to poultry spectra, reflecting differences in protein composition related to raw materials and hydrolysis enzymes (). Notably, the Fish hydrolysates dataset shows smaller standard deviations, indicating narrower variability in AMW compared to the wider span observed in the poultry datasets.

FIGURE 2

In Figure 3, mean second derivative FTIR spectra in the 1800–800 range are shown for each dataset (black for Fish, blue for Poultry 1, and green for Poultry 2), to highlight the band positions from literature which we refer to as set of biochemically motivated features (see Variable Selection Section 2.2). Pre-processing of these spectra by second derivative was performed for visualization purposes of the resolved chemical bands. It is observed that there are differences in the peak intensities at different band positions for different sample types, particularly at 1548 (amide II), 1515 (), 1452 (amide III), 1400 (). The band shape in amide I range (1650 ) differs across datasets, reflecting variations in protein and peptide structures. More information on the interpretation of the chemical spectral bands related to proteins in general and in hydrolysates can be found in the literature (; ; ; ; ; ). Peaks and shoulders within this region can be linked to differences in secondary structure and protein composition.

FIGURE 3

We start by establishing the benchmark for the calibration using the broadband raw spectra for calibration against the average molecular weight by Partial Least Squares Regression (PLSR). In the first three rows of Table 1 the PLSR results are reported for the independent test sets using: (i) the full range 4000–400 , (ii) the fingerprint 1800–800 , and (iii) the amide I range 1700–1600 . The regression results are reported using the coefficient of determination () and root mean squared error (RMSE) performance on the independent test set. For a PLSR analysis performed using the full spectral range (4000–400 ), the values obtained were 0.78 for the Fish, 0.85 for the Poultry 1, and 0.84 for the Poultry 2 dataset. For the Fish and Poultry 1 datasets, there is a noticeable improvement in when narrowing the spectral region from the full range (4000–400 ) to the fingerprint range (1800–800 ). Specifically, the increases from 0.78 to 0.84 for the Fish, and from 0.85 to 0.93 for the Poultry 1 dataset. When further restricting the spectral region to the amide I range (1700–1600 ), the values remain high, indicating a strong predictive ability of the amide I range. For the Poultry 2 dataset there is overall improvement from the broadband ( = 0.85) to the amide I range ( = 0.9), but there is an unexpected drop of performance when considering the fingerprint range ( = 0.7). This could be explained by a higher non-informative variability in the part of the fingerprint range which is outside the amide I range. This high variability may dominate in the model establishment and give inferior models.

TABLE 1

FishPoultry 1Poultry 2
4000–400 0.78 (99)0.85 (395)0.84 (854)
1800–800 0.84 (83)0.93 (260)0.7 (1192)
1700–1600 0.84 (84)0.84 (408)0.9 (668)
Set biochem. Motivated feat0.58 (136)0.74 (515)0.62 (1332)
Laser wavenumbers0.83 (87)0.75 (507)0.83 (883)
Random 4000–400 0.37 (165 )0.59 (645 )0.56 (1423 )
Random 1800–800 0.66 (121 )0.74 (512 )0.69 (1215 )
Random 1700–1600 0.79 (95 )0.75 (497 )0.84 (865 )

R2 (and RMSE in parenthesis in g/mol units) results for PLSR regression in prediction on Fish, Poultry 1, and Poultry 2 datasets, using i. full spectral range (4000–400 ), ii. fingerprint range (1800–800 ), iii. amide I range (1700–1600 ), iv. set of biochemically motivated features, v. existing laser wavenumbers, and vi. vii. viii. random selection on the previously specified ranges (average of 30 runs standard deviation).

When performing PLSR using (iv) the set of biochemically motivated features for estimation of AMW protein content, the prediction ability of the model is considerably lower than when using the entire spectral regions as described previously. The results obtained using the single biochemically motivated features are 0.58, 0.74, and 0.62 for the Fish, Poultry 1, and Poultry 2 datasets, respectively (iv row Table 1). It is important to note that the results obtained using these 9 wavenumbers (wns) selected from the whole fingerprint range are qualitatively considerably lower than using solely selected wavelengths from the amide I range.

PLSR regression was also tested using the wns in the I range for which the laser is currently configured as reported in Section 2.2 (). The results are = 0.83 for both Fish and Poultry 2 datasets, and = 0.75 for the Poultry 1 dataset.

Additionally, we compared these results with PLSR models generated by using randomly selected variables within the three different spectral ranges evaluated (see the last three rows in Table 1). Specifically, using 10 randomly selected variables within the predefined ranges (i, ii, iii), the results (reported as the average of 30 runs standard deviation) were worse compared to models using the corresponding full spectral ranges. However, the results obtained after random variable selection show a clear improvement in performance when transitioning from broader to narrower spectral ranges. Specifically, the values increase as follows: 0.37 0.66 0.79 for Fish; 0.59 0.74 0.75 for Poultry 1; 0.56 0.69 0.84 for Poultry 2, arriving to match the results obtained using the full range (4000–400 ) when random selection is performed on the amide I range (1700–1600 ) for Fish and Poultry 2 datasets. This makes sense as the narrower the range, the more representative are the 10 randomly selected variables that populate the specific range. Therefore, the amide I range, which is a very narrow range and which alone achieved very high prediction ability, is expected to perform best with a set of randomly selected variables where 10 randomly selected variables may perform similar as 10 equally spaced variables. Results that are equal to or lower than the values achieved by random selection for each spectral range (e.g., = 0.66, 0.74, and 0.69 for Fish, Poultry 1, and Poultry 2 datasets, respectively, when limited to the 1800–800 range) could be considered as being obtained by chance.

While the results reported above address the manual variable selection of spectral regions or sparse variables, the following sections report results on the performance of the different variable selection methods (SPLS, SPA, PVS, PVR, CovSel, g-CovSel, w-CovSel) for each dataset, separately. We will focus on the suitability of the different methods to serve as a design tool for laser wavelength selection.

3.2 Variable selection methods comparison

3.2.1 Fish dataset

For the design of the actual laser tunable in the range from 1700 to 1600 , only this region is relevant for variable selection. However, to keep the discussion general and to compare with other spectral regions for which such a tunable laser in principle could be built or for which a set of single wavelength emission lasers could be used, we consider different spectral regions in the complete mid-IR range from 4000–400 .

In Table 2, the selection performed by the different variable selection methods (performed on 4000–400 , 1800–800 , and 1700–1600 ranges), and corresponding results in terms of (coefficient of determination) and RMSE (root mean squared error) in predictions are shown for the Fish dataset.

TABLE 2

Variable selection method4000–400 1800–800 1700–1600
Selection (n)R2 (RMSE)Selection (n)R2 (RMSE)R2 (RMSE)
SPLS1571–16060.81 (92 g/mol)1587–15940.81 (91 g/mol)0.84 (86 g/mol)
1639–16741650–1658 (10)
1851–1886 (57)
SPA1656, 1675, 15870.75 (106 g/mol)1650, 1674, 16350.82 (90 g/mol)0.82 (88 g/mol)
1843 (4)1411, 1143, 1589
1454, 1043, 1774 (9)
PVS1270, 2534, 24570.65 (125 g/mol)1623, 1392, 9770.59 (136 g/mol)0.79 (98 g/mol)
3278, 1402, 38911286, 1664, 1758
1699, 3427, 16481170, 1481, 1602
624 (10)815 (10)
PVR1594, 1548, 17990.8 (94 g/mol)1594, 1548, 16870.84 (84 g/mol)0.86 (79 g/mol)
1697, 1421, 13961419, 1479, 1189
2765, 1839, 13031567, 1402, 1442
2349 (10)1027 (10)
CovSel1591, 1654, 18720.74 (107 g/mol)1591, 1654, 17990.84 (85 g/mol)0.86 (79 g/mol)
1675, 2204, 1043 (6)1675, 1043, 1149
635, 1415, 800
1496 (10)
G-CovSel 1604–1575 (16) 0.53 (145 g/mol) 1604–1575 (16) 0.53 (145 g/mol)0.42 (161 g/mol)
1670–1641 (16) 1670–1641 (16)
1934–1801 (70) 1799–1762 (20)
1683–1670 (8) 1683–1670 (8)
2316–2073 (127)All 0.87 (76 g/mol) 1053–1035 (10)All 0.82 (88 g/mol)
W-CovSel 1593–1589 (3) 0.33 (173 g/mol) 1593–1589 (3) 0.33 (173 g/mol)0.37 (167 g/mol)
3510–3506 (3) 1799–1795 (3)
1647–1643 (3) 1648–1644 (3)
1826–1822 (3)All 0.67 (122 g/mol) 1683–1679 (3)All 0.83 (86 g/mol)

Selected variables (and total number in parenthesis, out of a maximum of 1866, 519, and 52 variables) in order of selection, and (and RMSE in parenthesis) results for PLS regression in prediction on fish 95 samples dataset (66:29 train:test set), using the selection from different methods (SPLS, SPA, PVS, PVR, CovSel, g-CovSel, w-CovSel) applied within i. full spectral range (4000–400 ), ii. fingerprint range (1800–800 ), and iii. amide I range (1700–1600 ). The g- and w-CovSel results are for the first selected interval considered singularly , and for all the selected intervals (All) considered together.

The best result ( = 0.87) is obtained letting the g-CovSel method selecting a total of five groups of variables within the full range (4000–400 ) with correlation and covariance thresholds set to 0.95 and 0.7, respectively, as reported in Methods Section 2.2.4.2. The five selected groups are 1604–1575, 1670–1641, 1934–1801, 1683–1670, and 2316–2073 in order of selection comprising in a total of 237 variables. It can be seen that each group covers a relatively narrow spectral region. The best prediction is obtained when using all variables from all five groups.

In contrast to the g-CovSel method which selects groups of variables, the PVR and CovSel methods select discrete wavelengths based on the maximum correlation (PVR) and covariance (CovSel) values between predictive and response variables. For the PVR and CovSel method, very high predictive performance is obtained when selecting a maximum of 10 discrete wavelengths within the range 1700–1600 . This again indicated the strong predictive ability of the amide I range.

The w-CovSel method can be used to select narrow intervals akin to fine tuning around discrete laser wavelengths relevant for ICL laser (; ). We use it to select a maximum of five intervals made of 3 consecutive variables in each, according to maximum covariance between predictive and responsive variables. We observe, that four windows were selected for both the full (4000–400 ) and the fingerprint range (1800–800 ). The selected four windows for the two ranges are 1593–1589, 3510–3506, 1647–1643, and 1826–1822 for the full range, and 1593–1589, 1799–1795, 1648–1644, 1683–1679 for the fingerprint range in order of selection, for a total of 12 variables for each range. As can be seen, the first chosen interval is the same for the full and the fingerprint range.

The worst performance ( = 0.33) is obtained by using only the first interval alone (1589–1593 ). Using one interval with 3 consecutive and correlated variables is not expected to produce good prediction results. This is the same trend for all cases where only one interval of highly correlated variables are chosen as for the g- and w-CovSel (see Table 2). For the amide I range 1700–1600 , we report the results for the first group selected ( = 0.37) in Table 2. Other results for g- and w-CovSel can be found in the Supplementary Material. Selecting up to five narrow intervals within the amide I range by w-CovSel gives very good prediction ability, which is inline with previous results were we saw that three fine tuned single emission points obtained very good results for prediction of lipid profiles ().

3.2.2 Poultry 1 dataset

In Table 3, the variable selection performed by the different methods in the ranges 4000–400 , 1800–800 , and 1700–1600 is shown for the Poultry 1 dataset consisting of 133 samples.

TABLE 3

Variable selection method4000–400 1800–800 1700–1600
Selection (n)R2 (RMSE)Selection (n)R2 (RMSE)R2 (RMSE)
SPLS1668–1633 (19)0.47 (741 g/mol)1585–15810.85 (387 g/mol)0.45 (749 g/mol)
1660–1652
1745–1735 (13)
SPA1517, 1591, 16500.75 (510 g/mol)1654, 1558, 16250.81 (436 g/mol)0.82 (435 g/mol)
1195, 3220, 19671440, 1120, 1685
667 (7)1591, 1799 (8)
PVS3270, 1785, 23500.66 (593 g/mol)1546, 854, 13230.79 (463 g/mol)0.78 (478 g/mol)
1633, 2846, 37801627, 1780, 1081
852, 1388, 13551581, 1681, 1182
540 (10)997 (10)
PVR1203, 2960, 29890.93 (266 g/mol)1203, 1527, 10180.84 (401 g/mol)0.81 (446 g/mol)
1623, 3116, 17011627, 1778, 1012
1477, 3016, 19881392, 1311, 1274
1195 (10)1654 (10)
CovSel1656, 2512, 15190.8 (450 g/mol)1656, 1583, 17450.9 (316 g/mol)0.81 (446 g/mol)
1124, 3095, 7561124, 1529, 1799
3997, 1868, 39941053, 894, 1560
1585 (10)1620 (10)
G-CovSel 1670–1627 (23) 0.71 (543 g/mol) 1670–1627 (23) 0.71 (543 g/mol)0.71 (543 g/mol)
2632–2374 (135) 1799–1712 (46)
1539–1510 (16) 1535–1508 (15)
1137–1112 (14) 1135–1112 (13)
3213–2850 (14)All 0.87 (365 g/mol) 1614–1581 (18)All 0.82 (426 g/mol)
W-CovSel 1658–1654 (3) 0.32 (835 g/mol) 1658–1654 (3) 0.32 (835 g/mol)0.32 (835 g/mol)
2553–2549 (3) 1760–1756 (3)
1126–1122 (3) 1519–1515 (3)
1587–1583 (3) 1604–1600 (3)
609–605 (3)All 0.87 (365 g/mol) 950–946 (3)All 0.78 (488 g/mol)

Selected variables (and total number in parenthesis) in order of selection, and (and RMSE in parenthesis, out of a maximum of 1866, 519, and 52 variables) results for PLS regression in prediction on Poultry 1 dataset (93:40 train:test set), using the selection from different methods (SPLS, SPA, PVS, PVR, CovSel, g-CovSel, w-CovSel) applied within i. full spectral range (4000–400 ), ii. fingerprint range (1800–800 ), and iii. amide I range (1700–1600 ). The g- and w-CovSel results are for the first selected interval considered singularly , and for all the selected intervals (All) considered together.

The best predictive performance ( = 0.93) is obtained by using the 10 discrete variables selected by PVR, choosing variables according to the maximum correlation between the predictor and response variables. The selected wns are 1203, 2960, 2989, 1623, 3116, 1701, 1477, 3016, 1988, and 1195 in order of selection. This result is comparable to the result obtained using the full spectral range 1800–800 (see results in Table 1 for the Poultry 1 dataset).

Applying the w-CovSel method in order to choose a maximum of five intervals made of 3 consecutive wns each, five intervals are selected for both full (4000–400 ) and the fingerprint range (1800–800 ). The selected intervals are 1658–1654, 2553–2549, 1126–1122, 1587–1583, and 609–605 for the full and 1658–1654, 1760–1756, 1519–1515, 1604–1600, and 950–946 for the fingerprint range in order of selection. The five intervals comprise in total 15 variables. As before, the worst result ( = 0.32) is obtained using only the first selected interval (1658–1654 ) which is the same for the selection performed on both the full and the fingerprint range. However, we observe that when only using one selected group of variables identified by g-CovSel and comprising 23 variables within the region 1670–1627 that a highly predictive model is achieved with a = 0.71. The same group of variables is selected by g-CovSel for all three ranges. For more detailed results for variable selection by g-CovSel and w-CovSel see Supplementary Material.

For the SPLS variable selection we set the sparsity to select 99% in the full (4000–400 ) and fingerprint range (1800–800 ), and to select 90% of the variables in the amide I range (1700–1600 ). The prediction results are surprisingly low with = 0.47 and 0.45 when using the selection performed within the full, and amide I ranges respectively. However, for the fingerprint range the predictive ability for the variables selected by SPLS is very good. The selection on the full range includes a total of 19 variables between 1668 and 1633 , while when the selection is performed in the amide I range only 4 variables are selected (variables not shown/reported). For the fingerprint range a total of 13 variables was selected within different intervals spread out over this range. We observe that the SPLS selects from the full range almost the same variables as the g-CovSel selects as the first group of selected variable selected from all three ranges (see Table 3). However, the 4 additional variables selected by g-CovSel within the range 1670–1627 give a considerably better prediction ability, while SPLS selects within the 1668–1633 , where the predictive ability does not seem to be so strong. Decreasing the sparsity in SPLS to 80% on amide I and 90% on the full range lead to high prediction results which are comparable to the results obtained by g-CovSel (results not shown). As for the Fish dataset, more results for the g-CovSel and w-CovSel are reported in the Supplementary Material.

3.2.3 Poultry 2 dataset

In Table 4, the selection performed by the different variable selection methods (performed on 4000–400 , 1800–800 , and 1700–1600 ranges), and corresponding results in terms of and RMSE in predictions are shown for the Poultry 2 dataset.

TABLE 4

Variable selection method4000–400 1800–800 1700–1600
Selection (n)R2 (RMSE)Selection (n)R2 (RMSE)R2 (RMSE)
SPLS3456–34210.9 (690 g/mol)1799–17910.86 (799 g/mol)0.87 (775 g/mol)
1664–16291656–1648
1569–15661529–1521 (15)
1539–1510 (57)
SPA1527, 1660, 16000.86 (815 g/mol)1523, 1652, 11240.7 (1196 g/mol)0.86 (819 g/mol)
1822, 1174, 8521359, 1625, 1600
2586, 2177 (8)954, 1799 (8)
PVS3184, 1465, 35290.5 (1529 g/mol)1635, 948, 12990.63 (1313 g/mol)0.86 (805 g/mol)
2376, 1243, 6341498, 844, 1749
1629, 2869, 15441145, 1558, 1143
1101 (10)1440 (10)
PVR1654, 1643, 16450.88 (758 g/mol)1654, 1525, 17780.7 (1889 g/mol)0.86 (813 g/mol)
3270, 1691, 7051130, 1799, 935
3353, 3164, 24331413, 1149, 1591
1837 (10)898 (10)
CovSel1654, 3434, 15270.85 (853 g/mol)1654, 1523, 17990.79 (1002 g/mol)0.87 (769 g/mol)
1828, 1189, 23661122, 885, 1625
1600, 838 (8)1598, 1432 (8)
G-CovSel 1664–1629 (19) 0.88 (750 g/mol) 1664–1629 (19) 0.88 (750 g/mol)0.88 (750 g/mol)
3511–3375 (72) 1533–1515 (10)
1537–1510 (15) 1799–1762 (20)
1880–1787 (49) 1222–1101 (48)
1299–1159 (74)All 0.89 (706 g/mol) 840–800 (22)All 0.68 (1220 g/mol)
W-CovSel 1656–1652 (3) 0.84 (875 g/mol) 1656–1652 (3) 0.84 (875 g/mol)0.84 (875 g/mol)
3436–3432 (3) 1625–1621 (3)
1944–1940 (3) 1799–1795 (3)
1124–1120 (3) 1124–1120 (3)
1583–1579 (3)All 0.83 (901 g/mol) 806–802 (3)All 0.68 (1222 g/mol)

Selected variables (and total number in parenthesis, out of a maximum of 1866, 519, and 52 variables) in order of selection, and (and RMSE in parenthesis) results for PLS regression in prediction on Poultry 2 dataset (190:95 train:test set), using the selection from different methods (SPLS, SPA, PVS, PVR, CovSel, g-CovSel, w-CovSel) applied within i. full spectral range (4000–400 ), ii. fingerprint range (1800–800 ), and iii. amide I range (1700–1600 ). The g- and w-CovSel results are for the first selected interval considered singularly , and for all the selected intervals (All) considered together.

The best result ( = 0.9) is obtained when variable selection is performed by SPLS on the full range (4000–400 ) removing 99% of variables. SPLS selects more or less continuous ranges of variables, i.e. 3456-3421, 1664–1629, 1569–1566, and 1539–1510 . The results shows that in total 57 variables were selected, which means that much less than the 99% were eventually removed by SPLS. The same method applied to the fingerprint range (1800–800 ) provides = 0.86 with only 15 variables selected within 1799–1791, 1656–1648, and 1529–1521 . This shows again that the amide I has a very high prediction ability.

R2 = 0.88 is obtained when using the 10 discrete variables selected by PVR, which bases the selection on maximum values of correlation between predictive and responsive variables. PVR, as well as PVS, tend in general to select the maximum number of variables allowed, which is 10 also in this case (1654, 1643, 1645, 3270, 1691, 705, 3353, 3164, 2433, and 1837 in order of selection). The same result in terms of is obtained when using only one group among the five selected by g-CovSel method comprising 19 variables within the 1664–1629 region. The same group of variables is selected by g-CovSel for all three ranges.

For g- and w-CovSel methods, Table 4 shows that when variable selection is applied with these methods to the fingerprint range (1800–800 ), the model performance actually has a lower performance when all five selected groups or intervals are included compared to including only one group of variables. For g- and w-CovSel methods, it is in general expected that the prediction performance improves when increasing the number of selected intervals. For g-CovSel, the five selected groups were 1664–1629, 1533–1515, 1799–1762, 1222–1101, and 840–800 , corresponding to a total of 119 variables. For w-CovSel, the five selected intervals were 1656–1652, 1625–1621, 1799–1795, 1124–1120, and 806–802 , for a total of 15 variables. In both cases, the resulting performance was limited to = 0.68. By contrast, using only the first group or interval led to much better results. With g-CovSel, the first group (1664–1629 , 19 variables) alone yielded = 0.88. With w-CovSel, the first interval (1656–1652 , just 3 variables) achieved = 0.84. This is the only case where the first interval of w-CovSel considered alone is enough for high predictive performance (see previous Results Sections 3.2.1, 3.2.2).

In general, the results are all satisfactory except for the low predictive ability of the 10 single variables selected by the unsupervised method PVS, which is very similar to the PVR method, with the difference that the PVS method does not use a response variable in the selection process. The PVS method bases the selection only on the variance of the predictive data. The selections performed by the method on full (3184, 1465, 3529, 2376, 1243, 634, 1629, 2869, 1544, and 1101 in order of selection) and fingerprint (1523, 1652, 1124, 1498, 844, 1749, 1145, 1558, 1143, and 1440 in order of selection) results in predictive performances of = 0.5 and 0.63, respectively.

4 Discussion

4.1 Key spectral regions, informative variables, and existing laser

The selection of wavelength(s) for the design of a laser for protein characterization requires a precise understanding of which spectral regions provide the most informative and discriminative features for the target application. Broadband spectroscopic techniques, such as Fourier-Transform Infrared (FTIR) and Raman spectroscopy, collect data across a wide wavenumber range, much of which may contribute minimally or redundantly to estimate protein characteristics from the spectral features. For the development of miniaturized, application-specific devices, tunable lasers operating over narrow spectral regions with sparse, precisely selected wavelengths are highly desirable. To compare the variables selected by different variable selection techniques, we established the predictive power of continuous ranges (full 4000–400, fingerprint 1800–800, and amide I 1700–1600 ranges) in order to benchmark results. Normalization and baseline corrections are no straight forward on selected variables as they consume degrees of freedom from the already small set of selected variables. In addition, calculating of derivatives is not possible on discrete selected variables. Therefore, we chose to skip the pre-processing for the benchmark dataset to have comparable results for the benchmark data and the datasets with selected variables. This means no spectral pre-processing other than mean centering has been applied to obtain the regression results presented in Table 1. Hence, the results reported here show slightly lower predictive ability than in the studies (, ), where pre-processing was used to optimize prediction results.

In order to assess the predictive ability of a selection of laser wavelengths according to chemically and biologically meaningful bands in the mid-infrared spectral range, we evaluated as well a selection of wavelengths according to literature (, ). We used a set of nine single discrete wavenumbers (wns) that have been chosen with respect to their relevance for protein analysis in hydrolysates. The wns identified as important are: 1045 (collagen), 1081 (CO), 1400 ( terminal), 1452 (amide III), 1515 ( terminal), 1548 (amide II), 1630 (collagen), and 1645 (amide I); 1740 has been included as baseline. Using just the set of biochemically motivated features, provides suboptimal results as shown in row (iv) of Table 1. Based on the presented results, variable selection purely based on literature research looking for the best biochemically motivated features is not the best approach for laser building in the mid-infrared range, especially when limited in the actual number of data points to include, as in this case it has been set to have a maximum of nine.

The existing laser developed by the German company nanoplus and reported in (; ) consists of 19 data points within the 1700–1640 range. Among the 19 wns, it has been possible to use only 11, as reported in the Variable Selection Section 2.2. Using only these 11 wns for the PLSR predictions, the results are overall satisfying with 0.75. It is important to remember that the spectral and digital resolution of the available data does not have to be compared with a possible laser within this range, which would be able to reach higher resolutions.

We have demonstrated that selecting a sufficient number of data points (e.g., 10) within the appropriate range (1700–1600 ) can achieve an value greater than 0.75 across all datasets, even when the points are selected randomly (see the last row of Table 1). Comparable results can also be obtained using the existing laser wns. However, employing variable selection methods, such as the w-CovSel results for the amide I range, yields better outcomes than random selection procedures and outperforms the results obtained using the existing laser wns without further optimization (see Supplementary Table S2).

Notably, while the laser can theoretically emit at more than 1000 wns through temperature and current modulation (as discussed in the introduction), this optimization has not yet been fully explored. Currently, we can only confirm the actual emission for the wns tested by the laser provided by nanoplus. This indicates that future optimization of the emission could further enhance the laser’s prediction capabilities. In general, variable selection methods applied within 1800–800 range outperform the results obtained by choosing a set of wns relying only on the literature (which provide unsatisfactory 0.75 for all datasets).

Each method selects at least one wavenumber (or groups/intervals in case of g- and w-CovSel) within the amide I range (1600–1700 ). It is interesting to see, however, how some methods do not select any wns around 1630 or 1645/50, as literature would suggest, still providing good results (e.g., PVR for Fish - even when restricted to fingerprint range - and Poultry 1; SPA for Poultry 2).

Another important pair of wns suggested in the literature is 1400 and 1452–1. However, these wns are almost never selected, with a few exceptions (e.g., SPA, PVR, and PVS for Fish; SPA for Poultry 1; and PVS for Poultry 2). It is important to note that SPA, PVS, and PVR are all methods for the selection of single discrete variables. These methods are primarily based on variance and correlation measures. Consequently, if there is a strong and structured variation pattern in the data that is not correlated with AMW, these methods may still identify such a pattern as important, although it may not contribute significantly to the predictive ability of the variable set. Specifically, the absorption at 1400 corresponds to the terminal, which, according to the literature, is a critical wavenumber for assessing protein AMW. This is because it provides insight into the number of terminals present, which can serve as an indicator of peptide chain length. Nonetheless, the band at 1400 cm -1 does not improve the predictive ability of the model. This suggests that the variability associated with this wavenumber may not be particularly relevant for predicting AMW.

The amide I range is particularly informative, as its band positions are highly sensitive to protein composition, peptide length, and secondary structures such as -helices and -sheets. Previous studies have demonstrated its central role in determining protein structure using FTIR spectroscopy (; ; ). More recently, it has also been shown that the amide I provides key chemical information on protein composition and dominates regression models for predicting average molecular weights of proteins and peptides (). Even randomly selected wns from this range can capture these structural features, highlighting the central role of amide I analysis in protein characterization.

4.2 Variable selection methods comparison

G-CovSel is a method designed to select groups of variables, where each group corresponds to specific chemical contributions. The method operates by considering the correlation and covariance between predictive and response variables, as well as within the variables selected for each group. This makes g-CovSel particularly valuable for identifying groups of variables that capture the most significant underlying chemical information, especially for simple components (e.g., water content) (). Unlike its original implementation (CovSel), g-CovSel is not specifically designed to select a set of non-redundant and completely uncorrelated variables. As a result, the selected groups often consist of more or less continuous ranges, where often a large number of variables is selected (see Tables 24). The results indicate that g-CovSel does not gain additional information from any specific spectral range cut, even when the selection is restricted to very small ranges, such as 1700–1600 . However, this group-selection approach is well-suited for identifying “large” ranges or bands that can serve as a foundation for more specific single wns selection or for designing a “wide” tuning laser.

It is important to note that the groups of variables selected by g-CovSel are intended to be used in combination with one another (e.g., the second or third group alone is not expected to have high predictive ability). This is confirmed by the Fish dataset and the Poultry 1 dataset. When including the five groups selected by g-CovSel within the full (4000–400 ) and fingerprint (1800–800 ) range the predictions always result in 0.8 for Fish and Poultry 1 datasets. However, for the Poultry 2 dataset, a selection performed in the fingerprint range, including the five groups results in a decrease in the performance ( = 0.68). The same trend (and same ) is observed for the Poultry 2 dataset when performing the selection of five intervals made of 3 consecutive variables by w-CovSel on the fingerprint range.

When using the g-CovSel algorithm, it is not possible to directly control the number of variables included in each selected group. As a result, the total number of variables selected across the five groups is often quite high. For example, applying the method to the Fish dataset results in a total of 237 selected wns, with contributions from multiple groups: 16 variables from the first group, 16 from the second, 70 from the third, and so on (see Table 2). This would require incorporating over 200 wns into the laser design, which may pose practical challenges. To address this, one could consider using fewer than five groups or directly selecting a reduced number of groups, as g-CovSel often delivers satisfactory results with only two or three groups. For instance, for the Fish dataset, an 0.8 can be achieved by including just two groups (see Supplementary Table S1). This would significantly reduce the number of variables to include, e.g., using two groups instead of five for the Fish dataset reduces the total from 237 variables to just 32 (see Supplementary Table S1).

Using the w-CovSel algorithm, we selected up to five intervals, each consisting of 3 consecutive variables. The method demonstrates an improvement in prediction ability when restricting the selection range from the full (4000–400 ) to the fingerprint (1800–800 ) range. However, for the Poultry 1 dataset, this improvement is not observed. Even when considering all five selected intervals, the prediction ability is higher for the full range ( = 0.87) compared to the fingerprint range ( = 0.78). In such cases, one could consider increasing the number of allowed intervals. This approach, however, is not explored here due to practical considerations related to instrument design. A very large number of intervals (e.g., 20) is also undesirable for practical implementation. Additionally, the w-CovSel algorithm is particularly suited for selecting small windows of consecutive variables, making an increase in window size an unsuitable alternative.

The g-CovSel and w-CovSel algorithms are designed to improve predictive ability as the number of selected groups increases (here, from one to a maximum of five). For the Poultry 2 dataset, when the selection is restricted to the fingerprint range (see Table 4), an unexpected decrease in predictive performance is observed as the number of selected intervals increases. Specifically, for both variants of the CovSel method, the value drops from = 0.88 and = 0.84 for g-CovSel and w-CovSel, respectively, using only the first selected interval, to = 0.68 when five intervals are selected. CovSel-based algorithms rely on covariance, a combination of correlation and variance, which can result in the selection of variables with high variability rather than strong covariance, particularly in highly variable datasets like Poultry 2. This highlights how including uninformative variables, rather than carefully selecting the most relevant ones, can significantly impact predictive performance. A similar effect is observed in many other cases, particularly for w-CovSel, as shown in the (Supplementary Table S2). This underscores the importance of selecting the correct window size and region to optimize performance.

The variables selected by w-CovSel within the 1700–1600 range are detailed in the Supplementary Table S2, specifically in the first row (representing the first selected interval) of Supplementary Table S2. As part of an additional exploration of the w-CovSel method, larger window sizes were tested: 10 consecutive variables for the selection within the 1800–800 range, and both 10 and 50 for the 4000–400 range (as opposed to the default size of 3). Additionally, a larger number of intervals (five instead of a single interval) was tested for the 1700–1600 cm -1 range. The results of these explorations are provided in the Supplementary Table S2). The findings indicate that a window size of 3 consecutive variables is optimal for interval selection, balancing performance and relevance.

SPA, PVR, and CovSel are all methods designed to select a collection of single discrete variables, which in this study was set to 10. These methods exhibit mixed and non-constant behavior across different datasets. For the Fish dataset, the predictive ability increases consistently as the range dimension decreases. However, for the Poultry 1 dataset, predictive performance improves when narrowing the range from the full (4000–400 ) to the fingerprint (1800–800 ) range, but it declines when further restricted to the smaller amide I range (1700–1600 ). In contrast, for the Poultry 2 dataset, the predictive performance decreases when moving from the full spectrum to the fingerprint range but improves again when further narrowing down to the amide I range.

It is important to note that all the mentioned methods (SPA, PVR, and CovSel) provide an effective selection of single wns, achieving values greater than 0.8, especially when restricted to the fingerprint and amide I ranges. Notably, PVR and CovSel often yield satisfactory results even when the selection is performed over the full range. The ability to select single discrete wns is particularly relevant for the development of devices utilizing multiple single-emission wavelengths (e.g., using several lasers) rather than a single tunable laser (; ).

The SPLS algorithm allows users to set the degree of sparsity, which determines the percentage of variables to discard. However, the actual percentage of removed variables may sometimes be lower or higher, because the regression coefficients are not always exactly zero (see Methods Section 2.2.1). The selection process can include consecutive variables or isolated ones, depending on the values of the regression coefficients. By carefully focusing the selection on the appropriate initial region and choosing the optimal degree of sparsity, the method can be applied in two ways: (i) as a screening tool for large spectral regions (i.e., using a low degree of sparsity to select many variables), or (ii) as a more targeted approach for selecting potentially discrete variables within smaller ranges.

PVS performs satisfactorily only when the selection is restricted to the amide I range (1700–1600 ). However, its performance in this range is very similar to that of a random selection (see the last row of Table 1). This suggests that the method yields satisfactory results only when a pre-selection is made to focus on a range containing highly relevant information and variability. In contrast, when applied to the complete spectral range, PVS struggles to identify the information relevant to AMW. Instead, it may inadvertently detect other strong variation patterns that are not associated with AMW, leading to suboptimal performance. High variance does not guarantee high predictive power, as it may result from factors like noise or other structured variation that does not provide meaningful predictive information.

Other methods, such as LASSO, genetic algorithms, and variable importance in Random Forest, were considered, but ultimately excluded due to several drawbacks. For instance, these methods often select a large number of variables or exhibit instability in selection; stochastic algorithms like Random Forest do not produce consistent variable selections across different runs. Additionally, genetic algorithms are constrained by the number of initial variables, with a recommended maximum of fewer than 200 variables, which is impractical for high-dimensional spectroscopic datasets. Furthermore, some of these methods, such as genetic algorithms, are computationally intensive, making them less feasible for this study (; ; ).

5 Conclusion

Variable selection methods provide a systematic, data-driven approach to identifying the most relevant wavenumbers (wns). These techniques highlight the spectral regions most strongly correlated with the desired protein characteristics while eliminating irrelevant or noisy signals. By focusing the design on a small number of key wavelengths, the resulting laser system can achieve higher accuracy, robustness, and sensitivity, ultimately enhancing device performance.

In addition to improving model efficiency, variable selection enhances the interpretability of results by aligning selected regions with known molecular vibrations (e.g., amide I and II, side-chain vibrations). This supports biologically informed instrument design. Notably, relying solely on expert knowledge or literature-defined wns may overlook dataset-specific features or subtle variations. Variable selection integrates domain expertise with empirical evidence, ensuring that selected regions are both meaningful and optimized for the data at hand.

In this study, we evaluated a broad set of variable selection methods to determine their effectiveness in identifying the most relevant features for laser design. Methods such as SPLS, g-CovSel, and w-CovSel are particularly suitable for selecting either groups or intervals of consecutive wns, making them applicable for both narrowly and more broadly tunable lasers.

Conversely, methods like SPA, PVR, and CovSel are effective when a very sparse selection is required (typically fewer than 10 wavelengths). These methods not only enable compact laser design but are also valuable in pinpointing the most informative wavelengths within a limited spectral range, particularly when used in combination with broader interval-based approaches. The unsupervised method PVS is designed to select discrete variables as well, however, it struggles to identify relevant variables outside a small, chemically relevant range. It is better suited for initial dimensionality reduction and data exploration rather than selecting a small set of variables.

In summary, we find that g-CovSel is effective for identifying the most important ranges or regions in the spectrum. SPLS can also serve this purpose, particularly when avoiding the use of a maximum degree of sparsity (e.g., 90) over large ranges such as the full spectrum. Conversely, SPLS can be applied with a maximum degree of sparsity in smaller ranges for more specific variable selection. W-CovSel is optimal for selecting groups of consecutive wns, with a window size of 3 being the best choice. It is important to note that this approach is new and for the first time presented in the scientific literature. It is specifically designed to select sparse wavelengths with fine tuning around these wavelengths. Therefore, the method was called window(w)-CovSel.

Finally, SPA, PVR, and CovSel are suitable for selecting individual wavelengths. While limiting the selection to smaller regions (e.g., the fingerprint or amide I range) is ideal, both PVR and CovSel also perform well when applied to the full spectral range.

Importantly, differences in selection strategy can lead to substantial variation in predictive performance, even within the same spectral region. This highlights the critical impact of both the variable selection method and the exact spectral range chosen. Such differences are especially important in laser engineering, where technical constraints, such as achievable emission intensity, tuning resolution, or fabrication limitations, may prevent access to all target wavelengths at equal performance. Variable selection therefore serves as a crucial bridge between high-dimensional spectral data and the practical development of targeted laser systems.

Beyond mid-IR lasers, the same strategies may inform the design of other advanced optical systems, such as hyperspectral imaging, Raman, or atomic force microscopy (AFM) spectroscopy platforms, where spectral sparsity, tuning precision, and data interpretability are similarly critical.

Statements

Data availability statement

The spectral raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.

Ethics statement

This study utilizes FTIR spectra of protein hydrolysate samples derived from previous research conducted at Nofima AS (As, Norway) (, ). Ethical approval was not required for these studies, as the samples originated from salmon and poultry by-products obtained through routine food industrial processes.

Author contributions

MA: Data curation, Formal Analysis, Investigation, Software, Validation, Visualization, Writing – original draft, Writing – review and editing. FM: Methodology, Resources, Software, Supervision, Writing – review and editing, Writing – original draft. BK: Data curation, Validation, Writing – original draft, Writing – review and editing. ME: Writing – original draft, Writing – review and editing. PK: Writing – original draft, Writing – review and editing. KL: Methodology, Writing – review and editing, Writing – original draft. BZ: Writing – review and editing, Writing – original draft. NA: Data curation, Writing – review and editing, Writing – original draft. AB: Methodology, Writing – review and editing, Writing – original draft. VT: Methodology, Writing – review and editing, Writing – original draft. KT: Methodology, Writing – review and editing, Writing – original draft. VS: Writing – review and editing, Writing – original draft. AK: Conceptualization, Funding acquisition, Project administration, Resources, Supervision, Writing – review and editing, Writing – original draft.

Funding

The author(s) declared that financial support was received for this work and/or its publication. This work was partially performed within the PHOTONFOOD project, which received funding from the European Union’s Horizon 2020 research and innovation programme under grant agreement No. 101016444 and is part of the PHOTONICS PUBLIC PRIVATE PARTNERSHIP. Further funding was received from the Research Council of Norway (NFR) within the Centre for Research-Based Innovation in Digital Food Quality under grant agreement No. 309259 and within the collaborative Project to Meet Societal and Industry-related Challenges called Smart Sensing platform for protein and peptide analysis’ under grant agreement No. 353091.

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

The author KT declared that they were an editorial board member of Frontiers at the time of submission. This had no impact on the peer review process and the final decision.

Generative AI statement

The author(s) declared that generative AI was not used in the creation of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Supplementary material

The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fphot.2026.1696425/full#supplementary-material

References

Summary

Keywords

laser wavelength selection, machine learning, mid-infrared spectroscopy, protein spectroscopy, tunable laser, variable selection methods

Citation

Aledda M, Marini F, Kafle B, Erdem MC, Karki P, Liland KH, Zimmermann B, Afseth NK, Biancolillo A, Tafintseva V, Tøndel K, Shapaval V and Kohler A (2026) Wavelength selection for laser design in mid-infrared spectroscopy. Front. Photonics 7:1696425. doi: 10.3389/fphot.2026.1696425

Received

31 August 2025

Revised

13 January 2026

Accepted

19 January 2026

Published

18 February 2026

Volume

7 - 2026

Edited by

Georgios D. Barmparis, Foundation for Research and Technology Hellas, Greece

Reviewed by

Guangdong Zhou, Southwest University, China

Dip Das, University College London, United Kingdom

Updates

Copyright

*Correspondence: Miriam Aledda,

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics