ORIGINAL RESEARCH article

Front. Artif. Intell., 18 May 2026

Sec. Machine Learning and Artificial Intelligence

Volume 9 - 2026 | https://doi.org/10.3389/frai.2026.1794183

Deep learning classification of reproductive tissue from ultrasound: sex determination in red abalone (Haliotis rufescens)

  • 1. Department of Computer Science and Engineering, University of California, San Diego, San Diego, CA, United States

  • 2. Halıcıoǧlu Data Science Institute, University of California, San Diego, San Diego, CA, United States

  • 3. Department of Computer Science, University of California, Davis, Davis, CA, United States

  • 4. Bodega Marine Laboratory, University of California, Davis, Bodega Bay, CA, United States

  • 5. Department of Animal Science, University of California, Davis, Davis, CA, United States

Abstract

Introduction:

Accurate sex determination is critical for spawning success in both conservation breeding programs and commercial aquaculture, yet non-invasive methods remain limited in abalone species. Traditional approaches rely on visual inspection, which requires substrate detachment, can cause injury, and may induce premature gamete release. Here, we present the first application of machine learning to automate sex classification in red abalone (Haliotis rufescens) using non-invasive ultrasound imaging technology.

Methods:

We developed a labeled dataset of 246 high-quality ultrasound images from 44 individuals and benchmarked seven convolutional neural network architectures: VGG16, VGG19, ResNet50, ResNet101, YOLOv8, YOLOv11, and a custom convolutional neural network. Data partitioning by individual identity was essential to prevent artificially inflated accuracy from image leakage across splits.

Results:

The YOLOv8 architecture achieved the highest test accuracy of 85.7% (precision: 0.905 male, 0.816 female; recall: 0.845 male, 0.899 female), outperforming both classical architectures and custom models. Interestingly, a custom reverse VGG architecture with decreasing channel depth outperformed standard VGG models, suggesting that early channel compression may help combat ultrasound speckle noise.

Discussion:

Feature activation maps confirmed that models learned to attend to gonadal tissue rather than imaging artifacts. We also demonstrate inference on NVIDIA Jetson edge devices, enabling real-time classification suitable for field deployment. This framework establishes the feasibility of automated, non-lethal sex determination for mollusks and lays the groundwork for applications to endangered abalone species where traditional invasive methods are prohibited.

1 Introduction

1.1 Conservation background

In the 1970s, populations of California abalone (Haliotis spp.) began to decline due to historical overharvesting, disease, starvation, and other climate change factors (; ; Hobday and Tegner, 2000; Karpov et al., 2000; Hobday et al., 2000; Rogers-Bennett, Laura et al., 2010; Rogers-Bennett et al., 2021; Peters et al., 2024). Due to the dramatic declines in California abalone, all species are recognized as endangered or critically endangered by the International Union for Conservation of Nature (IUCN) (Peters and Rogers-Bennett, 2018, 2020, 2021a,b,c), whereas the black (H. cracherodii) and white abalone (H. sorenseni) are federally recognized as endangered species by the United States (Federal Register 66[103]; Federal Register 74[9]), with conservation becoming a critical tool for these species. As a result of these declines in abalone, there is currently no legal recreational or commercial abalone fishery along the west coast of the United States (Karpov et al., 2000). Restoration strategies have included the development of the Final Recovery Plan for Black Abalone (NMF, 2020) and the establishment of a White Abalone Captive Breeding Program (Rogers-Bennett et al., 2016). Recovery efforts seek to enhance abalone reproduction using a combination of approaches, including disease and parasite prevention, water quality and temperature management, and the use of non-invasive ultrasonography in conservation and production aquaculture to determine when animals are reproductively competent (Boles et al., 2022, ).

Abalone are herbivorous, dioecious (separate sexes) marine snails that release their gametes into the water column for external fertilization, and their gametogenic cycles vary by species, location, temperature, and season (Hahn, 1989). In captive breeding programs, abalone require regular reproductive monitoring to direct spawning activities. Traditionally, this method has relied on visual sexing of individuals, which can be unreliable due to changes in the gonad quality during the reproductive cycle. Traditionally, to visually inspect an abalone's gonad for sex identification, individuals must be detached from their substrate, which can induce the release of immature gametes due to handling stress and reduce spawning success (personal observation, S. Boles). Definitive sex determination may be performed using lethal histological examination, but this approach is prohibited for endangered species. To mitigate this stress response, ultrasound imaging technology has been developed as a non-invasive tool to visualize gonadal development and direct spawning activities without disrupting substrate attachment or initiating premature gamete release (Boles et al., 2022; Neylan et al., 2024).

A key component of successful captive propagation is the ability to accurately determine sex and reproductive condition prior to spawning. As broadcast spawners, abalone require coordinated gamete release for fertilization, and missed or mismatched spawning events not only reduce genetic yield but also waste limited hatchery time and resources. Traditional methods such as biopsy, dissection, or induced spawning are invasive, stress-inducing, and often prohibited for some endangered species. In farmed settings, these procedures can reduce animal welfare and operational efficiency as abalone lack external dimorphism, and sex determination via manual manipulation for gonadal assessment is difficult in immature animals. Further complicating sex determination, in non-reproductive abalone, reproductive status is not reliably size-specific. In red abalone, males may begin producing sperm at shell lengths as small as 75 mm, whereas females do not exhibit mature oocytes until approximately 105–130 mm, and individuals larger than 215 mm may exhibit reproductive senescence, including high rates of necrotic oocytes (Rogers-Bennett et al., 2003).

To address these challenges, ultrasound imaging technology has been validated as a non-lethal method to assess reproductive condition in abalone. This technique enables identification of gonadal vs. digestive tissue and provides a repeatable ordinal gonad index score in red abalone (Boles et al., 2022), which has since been applied to endangered abalone species, demonstrating the utility of ultrasonography in captive breeding programs (). While recent studies have proposed refined metrics, such as gonad area and relative average thickness, to improve interpretation (Zou et al., 2025), these methods still rely on manual measurements, are subject to operator bias, and remain labor-intensive. Even with standardized imaging protocols, expert interpretation remains a bottleneck for high-throughput applications.

1.2 Machine learning in aquaculture and biological imaging

Recent advances in machine learning (ML), particularly convolutional neural networks (CNNs), have enabled automated classification of complex biological structures from imaging data, including ultrasound. CNNs extract hierarchical features directly from raw images, bypassing the need for handcrafted descriptors and enabling high-accuracy classification even in noisy, real-world datasets. In medical imaging, CNN-based systems have achieved human-level performance on tasks ranging from tumor detection to retinal disease diagnosis (; Gulshan et al., 2016). In aquaculture settings, CNN-based models have been applied to detect disease symptoms, estimate biomass, and classify sex and ovarian stage in finfish using ultrasonography (; Graham et al., 2022). However, applications of ML to marine invertebrates remain rare, limited by anatomical complexity and the scarcity of labeled datasets ().

Abalone present a compelling model for extending these techniques to mollusks. Unlike most gastropods, whose reproductive organs are tucked deep within a spiral shell, the entire internal body of the abalone exists beneath the shell. The gonad/digestive gland complex is accessible to ultrasound imaging from the ventral foot surface. In large red abalone, this allows the entire gonad to be resolved within a single field of view, whereas in smaller individuals the entire soft body beneath the shell can be visualized, making abalone particularly well suited for non-invasive ultrasonography. Prior ML-based sex prediction studies in abalone have used the UCI Abalone Dataset (Nash et al., 1994), a benchmark dataset of physical measurements from Haliotis rubra that includes both non-destructive measurements (shell length, diameter, height, whole weight) and measurements that require destructive processing (shucked weight, viscera weight, dried shell weight, and growth ring counts obtained by sectioning the shell) (). Sex classification on this dataset has proven notably difficult, with reported accuracies typically falling in the 50–55% range and the best reported non-invasive result reaching only 56.9% after extensive polynomial feature engineering (). While these studies demonstrated that ML classifiers can differentiate sex from morphometric features at rates modestly above chance, the inclusion of destructive measurements in the feature set makes the approach incompatible with conservation goals where animals cannot be sacrificed, and the reliance on tabular morphometric data does not leverage the spatial information available in imaging modalities such as ultrasonography.

A practical constraint shared across medical and veterinary ultrasound domains is the limited availability of labeled training data. A scoping review of deep learning classification across medical imaging modalities found that the majority of ultrasound datasets were private, with studies spanning a wide range of dataset sizes including datasets below 1,000 images (Laçi et al., 2025). Lawley et al. (2024) demonstrated that neural network classification of 16 abdominal ultrasound cross-sections from a balanced subset of only 800 images achieved 79.5% accuracy, within 4.4 percentage points of the 83.9% attained by more complex architectures trained on the full dataset of 26,294 images; notably, the authors concluded that dataset size was a more important factor in determining accuracy than network selection. Similarly, Graham et al. (2022) trained CNN models for ovarian development classification in channel catfish (Ictalurus punctatus) using 931 ultrasound images and achieved median accuracies exceeding 98% for a binary classification task. These results demonstrate that useful classification accuracy is achievable from small ultrasound datasets when appropriate augmentation, regularization, and data partitioning strategies are employed, though larger datasets consistently improve performance (Lawley et al., 2024). This finding is consistent with our own experience in the present study, and we identify dataset expansion as a priority for future work (Section 4.6). An additional challenge common to ultrasound datasets is image quality variability; recent work has demonstrated that deep learning-based quality assessment can identify artifact-laden or non-informative frames prior to classification (; Kucharski et al., 2022), an approach we consider a natural extension of the pipeline presented here.

1.3 Study objectives

In this study, we present the first demonstration of ML-assisted sex classification in abalone using ultrasound imaging technology. We introduce a fully labeled dataset of red abalone (Haliotis rufescens) ultrasound images, a critically endangered species (Peters et al., 2021), and benchmark a suite of CNN architectures; VGG16, VGG19 (Simonyan and Zisserman, 2015), ResNet50, ResNet101 (He et al., 2016), YOLOv8, YOLOv11 (Jocher et al., 2023; Jocher and Qiu, 2024), and custom CNNs trained on confirmed male and female individuals. While our long-term goal is to assess reproductive maturity, the objectives of this study are to establish a high-throughput, non-invasive sexing tool applicable to both conservation and commercial hatchery workflows. Model performance is evaluated using accuracy, precision, and recall, with additional analyses on the effects of animal size and image quality. By integrating ultrasonography with automated image analysis, this work addresses a critical bottleneck in molluscan aquaculture and lays the foundation for future systems capable of linking soft tissue morphology with reproductive outcomes and spawning success.

2 Materials and methods

2.1 Abalone husbandry and sexing

Red abalone of two different size classes were sourced from The Cultured Abalone Farm (Goleta, CA). Individual animals were housed at the University of California, Davis Bodega Marine Laboratory in ambient oxygenated seawater flow-through conditions in a 9 L clear, polycarbonate tank and fed ad libitum a diet of Dulse (Devaleraea mollis) and ABKelp® (AlgalMar, Baja California, MEX). Red abalone were divided into two size classes: the small class (n = 41 individuals), which had a mean weight of 80.6 g (SD ± 15.1) and a mean shell length of 77.2 mm (SD ± 3.51), and the large class (n = 34 individuals), which had a mean weight of 259.7 g (SD ± 56.4) and a mean shell length of 110.53 mm (SD ± 15.3). Sex was determined by visual inspection of gonad coloration by The Cultured Abalone Farm management prior to shipment to the UC Davis Bodega Marine Laboratory. In reproductive red abalone, male gonads display a creamy white to tan coloration, whereas female gonads display a dark green coloration (Rogers-Bennett et al., 2004; Hahn, 1989).

2.2 Abalone ultrasound imaging protocol

A SonoSite Edge II Ultrasound System (FUJIFILM SonoSite Bothell, Washington) equipped with a HFL50 15-6-MHz transducer probe (Exam Type = Breast; Mechanical Index = 0.7; Read Depth = 6; Thermal Index = 0.1) was used to image red abalone gonad reproductive state. Non-lethal ultrasound examinations were performed by submerging abalone in seawater either attached to clear transparency copier film (3M #PP2950, Austin, Texas) or they remained affixed to their plastic housing (Boles et al., 2022). Red abalone of known sex were examined weekly using ultrasound imaging to characterize sex-specific gonadal architecture, supplemented by analyses of archived images from the same individuals and used as ground truth for training (Figure 1).

Figure 1

2.3 Data preparation and augmentation

Ultrasound images of red abalone were preprocessed through an automated pipeline implemented in Python v3.11 (Python Software Foundation, 2023) using Pillow v10.2.0 (), OpenCV v4.10.0.84 (), and NumPy v2.2.2 (Harris et al., 2020). The pipeline consisted of four sequential stages: border removal, region-of-interest (ROI) cropping, filename standardization and organization, zero-padding to uniform dimensions, and dataset partitioning.

2.3.1 Image cropping and organization

Images were first cropped using the Python Pillow library to remove borders and system overlays, which contained metadata including individual identifiers, sex labels, and ultrasound acquisition parameters. Retaining these overlays would risk models learning to classify sex from text rather than from gonadal morphology. Cropped images retained only the raw ultrasonography region. Images were then converted to grayscale and binarized using a fixed intensity threshold of 5 (on a 0–255 scale) to separate the abalone specimen from the dark imaging background. External contours were detected on the binary mask using cv2.findContours with the RETR_EXTERNAL retrieval mode, which identifies only outermost boundaries, and the CHAIN_APPROX_NONE approximation method, which preserves all contour points without simplification. The contour enclosing the largest area was selected as the specimen ROI, and its axis-aligned bounding rectangle was extracted via cv2.boundingRect. The original image was then cropped to this bounding rectangle, yielding specimen-only images of variable dimensions. This approach required no manual annotation; the low threshold value effectively distinguished the specimen from the uniformly dark ultrasound background across all images.

Once cropped, we organized images based on individual (denoted by unique identifiers), sex (Male [M] or Female [F]), and size (small or large). In order to reduce bias in our models we split our data by size, in case there were significant differences in the ultrasound images among the two sizes (Table 1). Unique individuals were identified based on data encoded in the filename string which contained their individual number (ID), location, sex, size, and date. Unique individuals were grouped by their ID, location, sex and size. Images were then deposited into two separate folders (small or large), followed by folders indicating their sex (M or F).

Table 1

LargeSmallUnknownTotal
F201 (22)263 (32)88 (24)552 (78)
M203 (25)241 (28)131 (40)575 (93)
Total404 (47)504 (60)219 (64)1127 (171)

Distribution of all unfiltered red abalone (Haliotis rufescens) images by sex and size class.

Unique individual counts are shown in parentheses.

“Unknown” refers to individuals whose size class was not recorded at the time of imaging.

2.3.2 Image quality filtering

Large abalone images were filtered for quality through visual inspection by E.A.S., who had no prior training in abalone ultrasonography. Images were excluded if they met any of the following criteria: (a) severe echoing artifacts in which the ultrasound beam reflected off the shell, creating duplicate or ghost structures; (b) absence of identifiable gonad and digestive gland anatomy; or (c) incomplete capture of the animal such that the field of view did not encompass the gonad region. Examples of excluded and retained images are shown in Figures 2, 3. Only images with obvious defects were removed; borderline cases were retained. This conservative threshold was chosen under the assumption that an operator with ultrasonography experience would filter more aggressively than an operator with none, and we therefore report results on a dataset that likely includes some suboptimal images. This filtering yielded a total of 246 high quality images from large individuals (Table 2).

Figure 2

Figure 3

Table 2

UnfilteredFilteredAugmented
F201 (22)114 (22)174 (22)
M203 (25)132 (22)199 (22)
Total404 (47)246 (44)373 (44)

Distribution of large red abalone (Haliotis rufescens) ultrasound images by sex at each stage of preprocessing: unfiltered, quality-filtered, and augmented.

Unique individual counts are shown in parentheses.

2.3.3 Dataset partitioning

The final dataset consisted of only large individuals, from which 4 individuals of each sex were randomly assigned to the validation and test sets, respectively. The remaining 28 individuals (14 of each sex) were allocated to the training set. All images from a given individual were confined to a single partition, ensuring that no specimen appeared across multiple splits. Because the number of images per individual was uneven, the resulting proportion of total images in each partition did not correspond exactly to a 90/10 or 80/20 ratio; however, the individual-level constraint was maintained throughout. Best models were selected based on validation metrics, and the held-out test set remained untouched until final evaluation. No models were selected based on test set performance.

An initial stage of data augmentation was performed to generate synthetic data for balanced individual representation in the training split. In creating the synthetic data, existing images were rotated up to 15 degrees in either direction, random brightness and contrast values were jittered by 10%, and arbitrary zooming up to 20% (Figure 3) was applied with the torchvision 0.17.2 library (Paszke et al., 2019).

We note that an earlier iteration of the pipeline used image-level random subsampling on an 80/20 train/validation split without individual-based constraints. This approach was abandoned prior to any reported experiments due to data leakage concerns, and all results presented in this manuscript reflect the individual-based partitioning scheme described above.

2.4 Logistic regression on flattened abalone images

To classify abalone sex (female vs. male) from ultrasound images, we employed a logistic regression framework applied to flattened image representations. All analyses were implemented in Python v3.10 (Python Software Foundation, 2023).

Images were loaded and preprocessed using Pillow v10.2.0 (). Each image was converted to grayscale and resized to 64 × 64 pixels, yielding a 4,096-dimensional feature vector per sample after flattening. The dataset partitions described above were maintained throughout all experiments, yielding training (n = 153 images; 80 F, 73 M), validation (n = 51 images; 22 F, 29 M), and held-out test (n = 42 images; 12 F, 30 M) splits. The corresponding individual counts for each split are reported in Table 3.

Table 3

SmallLarge
TrainValTestTrainValTest
F201+121 (22)40 (5)39 (5)80+60 (14)22 (4)12 (4)
M170+96 (20)30 (4)29 (4)73+67 (14)29 (4)30 (4)
Total371+217 (42)70 (9)68 (9)153+127 (28)51 (8)42 (8)

Distribution of red abalone (Haliotis rufescens) ultrasound images by sex across training, validation, and held-out test splits for small and large size classes.

Cell values are formatted as raw_count+augmented_count(unique_individual_count), where the first value is the number of original images, the second is the number of synthetic augmented images, and the parenthetical is the number of unique individuals.

Pixel intensities were standardized to zero mean and unit variance using feature-wise statistics computed exclusively from the training set. Standardization and all subsequent modeling steps were performed with scikit-learn v1.7.2 (Pedregosa et al., 2011b). To reduce dimensionality and mitigate overfitting, principal component analysis (PCA) was applied to the standardized features. The number of retained PCA components was treated as a tunable hyperparameter, with candidate values of 5, 10, 20, 30, 50, 75, and 100. Array operations were handled with NumPy v2.2.2 (Harris et al., 2020).

Hyperparameter selection was performed via grid search over the regularization strength C (0.001, 0.01, 0.1, 0.5, 1.0, 3.0, 5.0, 7.0, 10.0, 15.0, 20.0, 50.0, 100.0), penalty type (L1, L2), solver (liblinear, saga), and the number of PCA components listed above. To account for class imbalance across the splits, inverse class-frequency weighting was applied during training via the class_weight=~balanced~ parameter in scikit-learn, which scales the loss contribution of each class in proportion to its representation in the training set. Model selection was guided by balanced accuracy on the validation set, which averages per-class recall and thereby prevents the selection of degenerate classifiers that predict only the majority class.

The best-performing configuration was serialized to disk using joblib v1.5.2 (Joblib Development Team, 2024). Final evaluation was then conducted on the held-out test set, which remained untouched during all tuning and model selection steps. We report accuracy, per-class precision, per-class recall, and per-class F1 score for each of the training, validation, and test partitions. A confusion matrix and cumulative PCA explained variance plot were generated using matplotlib v3.10.7 (Hunter, 2007).

This procedure was conducted independently on two versions of the training data: the original (non-augmented) images and an augmented variant. The augmented dataset was prepared separately and subjected to the identical training, validation, and test pipeline described above, preserving the same held-out test partition across both experiments to enable direct comparison.

2.5 CNN model architectures

Training was done using various classification models implemented in PyTorch v2.2.2 (Paszke et al., 2019). These models included VGG16, VGG19 (Simonyan and Zisserman, 2015), ResNet50, ResNet101 (He et al., 2016), and our own custom-built CNN models which inverted the structure of the Convolutional 2D layers and the structure of the 1D layers of the VGG architecture. YOLOv8 and YOLOv11 models were trained using the Ultralytics library v8.3.3 (Jocher et al., 2023; Jocher and Qiu, 2024), which provides a PyTorch-based training pipeline. All training was performed on NVIDIA A100 GPUs via Expanse at the San Diego Supercomputer Center and Bridges-2 at the Pittsburgh Supercomputing Center. Additional data augmentation on top of the original synthetic data generation was performed according to YOLO defaults for all models, which included random horizontal flipping with probability 0.5, color jittering with brightness, contrast, saturation, and hue variations of 40%, 70%, 70%, and hue factor of 0.015, respectively. For each image, a rectangle region was erased with probability 0.4 and size as the defaults, and a random translation of 10% of the image size was applied. Training was performed using binary categorical labels. Their respective output layers were set to a single sigmoid with the BinaryCrossEntropy loss function. Batch sizes were set to 48, 32, 16 and 128 for custom VGG, ResNet, Custom Reverse, and YOLO models respectively. Model checkpointing and early stopping was implemented using the validation loss (0.70 minimum) for the models respective loss function with initial monitoring after 100 epochs and a patience of 500 epochs.

Models were initially trained on small animals using the same procedures described above, but due to a lackluster performance, this approach was abandoned in favor of transfer learning, in which each architecture was initialized with the weights from its best-performing large-animal model and training was continued on the small animal image dataset (Table 3). The same individual-based partitioning, augmentation, and early stopping procedures were applied to the small animal splits.

2.6 Hyperparameter tuning

Custom CNN models were hyperparameters tuned across various learning rates, dropout rates, layer depth, layer breadth and optimizers. We applied various degrees of dropout in order to prevent overfitting. In total 22 model structures were trained on with varying learning rates, dropout rates and optimizers for a total of 44 different parameter combinations for each of the model structures, for a total of 968 custom CNN models.

2.7 Evaluation metrics

Models were evaluated based on training and validation losses as well as precision, recall, and accuracy. Models were chosen based on validation accuracy, and finally tested using our test dataset. We report only the best models, based on validation recall, precision, and accuracy, for each of the various architectures (Table 4).

Table 4

ModelM Prec.M Rec.F Prec.F Rec.Train Acc.Val Acc.Test Acc.
YOLOv80.9050.8450.8160.8990.9610.8820.857
YOLOv110.5240.9520.9170.6670.9960.8630.738
ResNet-500.6670.6150.8330.8620.8110.8240.786
ResNet-1010.5000.6000.8670.8120.9570.8430.762
VGG-160.4170.2780.5670.7080.9570.7060.524
VGG-19-------
Custom-S0.9170.4400.5330.9410.9250.6470.643
Custom-M0.5000.3750.6670.7690.9320.7450.619
Custom-L0.7500.5620.7670.8850.9680.7450.762
Reverse-S0.4170.4170.7670.7670.9000.6860.667
Reverse-M0.6670.5710.8000.8570.9540.7060.762
Reverse-L0.9170.5790.7330.9570.9930.7840.786

Best model performance by architecture for sex classification of large red abalone (Haliotis rufescens) ultrasound images, selected on validation metrics.

Precision and recall are reported per class (M = male, F = female). Best values are shown in bold and second best are underlined. VGG-19 failed to converge to nontrivial results across all runs.

2.8 Feature extraction and visualization

We extracted the top five PCA components from images using the scikit-learn libraries in Python (Pedregosa et al., 2011a), and chose the top two components for K-means clustering. The optimal K clusters were determined using the elbow method using Within-Cluster Sum of Squares and the silhouette score. We also looked at plotting the third principal component against the first and second, as well as all top three using 3D matplotlib plotting libraries. The t-SNE dimensionality reduction technique (Maaten and Hinton, 2008) was used to visualize our model features using scikit-learn v1.5.0. Feature extraction was performed by extracting the weights of each of the best models and applying them to randomly subsampled images. Images were plotted using the OpenCV package.

2.9 EdgeAI device deployment

Using only our best models for edge device deployment, we present our best extra large YOLOv8 model. We exported the YOLOv8 model to ONNX format through Ultralytics export API which invokes torch.onnx.export. The ONNX model was then converted to a TensorRT engine using Ultralytics' onnx2engine utility with FP16 precision. The TensorRT engine was executed on a 8GB Jetson Orin Nano. Inference was made using a stream of prefiltered train and test images located in separate folders on the device.

3 Results

3.1 Data characteristics

Structure in the data was evaluated using filename-derived naming conventions, which allowed us to group images by individual identity, size, sex, and location. When individual identity was not considered and images were randomly split into training, validation, and test sets at the image level, the same animal could be represented in multiple sets, leading to biased training and overfitting. We found accuracy to be artificially inflated in images split without taking individuals into account. Once we enforced splitting by individual, model evaluation metrics were more in line with realistic performance, with initial validation and test accuracy at approximately 60%. These analyses indicated that a higher proportion of images from small animals were misclassified than those from large animals. We decided to split the data into large and small size classes, and initial training and validation results for large animals improved with metrics reaching approximately 70%, whereas models trained on small individuals remained near 55%. We continued to explore outliers and observed that images that were not clear, had imaging artifacts, or did not show clear gonads and digestive glands were often misclassified across various models. We then labeled them as poor-quality images, and manually removed them from training, we provide an example in Figure 2. This led to improved model performance and metrics.

3.1.1 K-means clustering

We applied K-means clustering to CNN-extracted ultrasound images features after reducing dimensionality using PCA, retaining the top three components. Cluster numbers from K = 2 to 9 were tested. The elbow method and silhouette analysis both indicated K = 2 as the optimal number of clusters (Supplementary Figure S1). Our PCA scatter plots (Supplementary Figure S2) showed overlap among large and small males and females, with no clear separation by sex or size.

3.2 Model performance comparison

As a baseline, we trained a logistic regression classifier on PCA-reduced flattened image features using the same individual-based data splits (Section 2.4). Without data augmentation, this resulted in 100%, 51%, and 71% accuracy in train, validation and test image sets, respectively, with the train-to-test gap indicating clear overfitting to the low-dimensional feature space. With data augmentation, accuracy was 67%, 51%, and 76% in train, validation and test image sets, respectively, partially mitigating overfitting but remaining near chance on the validation set. These results are consistent with the PCA scatter plots (Supplementary Figure S3), which show substantial class overlap. These results establish a non-deep-learning baseline against which all CNN architectures in Table 4 should be compared.

We then benchmarked CNN architectures, beginning with VGG and ResNet, followed by our reversed custom CNN. Our implementations of VGG and ResNet models resulted in the models quickly over-fitting. Dropout was added to our VGG and custom CNN models to prevent overfitting.

Our VGG16 and VGG19 models are identical to their original implementation with the exception of adding dropout. We performed hyper-parameter tuning across optimization functions, learning rates, batch size, activation functions and dropout rates. This resulted in 95.7%, 70.6% and 52.4% accuracy in train, validation and test image sets, respectively.

During systematic exploration of architectural variants, we evaluated an inverted VGG model in which the layer channel progression was reversed, starting with 512 channels and progressively reducing to 64, in both the 2D and 1D layer blocks. Regularization via dropout layers was implemented after each block. This resulted in better performance than the vanilla VGG models and was on par with some of our ResNet models, resulting in 99.3%, 78.4% and 78.6% (Reverse Large) accuracy in train, validation and test image sets, respectively. Our ResNet models performed better than our custom VGG, and custom CNN models. Our Resnet 101 achieved 95.7%, 84.3% and 76.2% accuracy in train, validation and test image sets, respectively. Our final implementation of the image classification models, YOLOv8x-cls (Jocher et al., 2023), resulted in 96.1%, 88.2% and 85.7% (X-Large) accuracy in the train, validation, and test sets, respectively (Figure 4). We also tested a newer implementation of YOLO11x-cls which resulted in 99.6%, 86.3%, and 73.8% accuracy, respectively.

Figure 4

3.3 Feature analysis

The t-SNE results from the YOLOv8 model illustrate some separation of male and female large abalone, as well as some that do not cluster together (Figure 5). We performed various feature extraction maps and visualized them (Figures 69). We can observe multiple activations across the gonad tissue for most of our best models, as well as activation in adjacent anatomical areas. As we progress through the layers, we continue to see activation of the gonad and a deactivation of the digestive gland, until the feature maps become too abstracted to interpret.

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

3.4 EdgeAI device inference

Inference times averaged 14.50ms per frame, but also required pre-processing and post-processing for a total of 1.35ms on the Jetson Orin Nano and a test accuracy of 85.7%.

4 Discussion

4.1 Principal findings

In this study, we present, to our knowledge, the first application of machine learning to automate sex classification in abalone using non-invasive ultrasound imaging technology. By developing an ultrasound image dataset for red abalone and benchmarking a logistic regression baseline alongside a suite of convolutional neural network architectures, we show that CNN models such as YOLOv8 and ResNet can distinguish between male and female gonadal morphology from soft tissue ultrasound images for this species, whereas linear classifiers on PCA-reduced flattened image features cannot. These results indicate that ultrasound based image classification in combination with machine learning technology can provide a practical, non-lethal tool for sex determination in both conservation and production aquaculture settings.

Our results highlight several practical insights for future deployment: (1) accurate model evaluation requires data partitioning by individual rather than by image to avoid artificially inflated performance; (2) image quality strongly influences classification accuracy, with low-contrast or artifact-laden scans contributing disproportionately to errors; and (3) size-dependent anatomical variation can bias model predictions, necessitating size-aware sampling or stratification. Although overall performance is promising, particularly for large individuals and high-quality images, the variability in model confidence underscores the need for continued dataset expansion, improved standardization of imaging protocols, and integration of automated quality control.

We initially explored various image filters such as non-local means denoising which is meant to improve ultrasound details by removing white speckle noise using different methods such as Gaussian blurs, but ultimately did not use these methods for our final models, as it did not improve model performance. Initial data preprocessing yielded deceptively good results, but the similarity of ultrasound images belonging to the same individuals meant that the naive splitting by image caused data leakages throughout the dataset, so the final preprocessing allocated splits by individuals instead. Additionally, the quality of small abalone ultrasounds were far inferior to those of the large abalone, and due to the nature of the small dataset at hand, we settled in favor of high-quality images only and relied on heavy data augmentation and dropout to prevent the overfitting of our models.

4.2 Architecture insights

The removal of small abalone images and the final filtering of low-quality large abalone ultrasounds raised model performance from near-random guessing accuracy of 60% toward the 80% mark, a significant performance increase. Our implementation of VGG models, despite strong data augmentation with YOLO presets, generally either failed to converge or performed worse than our YOLO models. We propose that due to the model age and lack of features present in models like ResNet with its skip connections or the lightweight and modular design of YOLO, training with VGG was unstable and inefficient due to the large parameter count. Since the VGG family are particularly large models fitted with the classification task of a tiny training set of 280 images with few options for regularization, it fell behind the other models consistently. This was also most likely due to how the VGG architecture starts with a low number of features but then expands into a larger number of features, doubling at each layer block (64 → 128 → 256 → 512 → 512). We reasoned that this expansion pattern, while effective for natural images, may be suboptimal for ultrasound data.

In standard VGG, early layers with few channels are tasked with extracting low-level features such as edges and textures, which then serve as building blocks for increasingly abstract representations in deeper layers. However, ultrasound images contain pervasive speckle noise that manifests as high-frequency texture patterns across the image. When VGG's early layers, constrained to only 64 channels, attempt to encode these patterns, they may allocate limited representational capacity to noise artifacts rather than biologically meaningful features. These noisy low-level representations then propagate through the network, potentially corrupting the higher-level features built upon them. In contrast, our reverse architecture begins with a large number of channels (512), allowing the model to capture a diverse set of features before progressively compressing them (512 → 256 → 128 → 64), which we propose forces the network to retain only the most discriminative information while discarding speckle-related noise. This design philosophy aligns with the Information Bottleneck principle (Tishby and Zaslavsky, 2015), which posits that optimal representations compress input information while preserving task-relevant features, precisely what our architecture enforces structurally through simultaneous channel and spatial reduction.

In response to the poor performance of the VGG model family, we additionally implemented various architectural changes as spin-offs of the VGG architecture, only modifying the depth and breadth of the layers, but performance remained similar to that of the VGG models. We also tested the Reverse-S/M/L family of models, obtained by literally reversing the convolutional layers of the VGG architecture and fine tuning the depth and breadth of each layer, starting with a large channel count and gradually reducing it over the layers (512 → 256 → 128 → 64). Counterintuitively, the models performed better than the VGG, and modified VGG models. CNNs are generally designed with inductive-biases toward detecting low level features first, gradually connecting them with long-range dependencies later on in the model. However, with each variant of the reverse models outperforming its forward counterparts by approximately 5% on average, it is very likely that this architectural change helps in the abalone ultrasound domain. Notably, recent work by (Liu et al. 2025) systematically evaluated filter placement topologies in ResNet-based architectures and introduced “RevNet” (Reverse ResNet), which implements the same decreasing filter configuration (512 → 256 → 128 → 64). Their findings demonstrated that this contrarian strategy, reducing filters by half across successive layers, can achieve performance on par with, or superior to, the conventional approach on standard benchmarks. Our results extend this finding to the ultrasound imaging domain, representing, to our knowledge, the first application of this inverted channel topology to marine invertebrate tissue classification.

First, the grainy noise in ultrasound imaging hurts conventional CNN models by introducing a large number of insignificant local artifacts. These small-scale details, such as speckle patterns, can dominate the low-level feature extraction process. On the other hand, the reverse model starts by encoding a wide variety of features, and the gradual reduction of channels over time forces the model to compress the information and wipe out unmeaningful small edge details. Second, compression of channels in the reverse model aids in reducing overfitting, particularly in settings with limited data. With a small dataset, it is very easy for standard CNNs to learn to rely on minute details like speckles in the ultrasound data, which are actually meaningless patterns. We see this in the failure of convergence in the VGG-19, which is a particularly large and monolithic model that significantly overfits on the training data. On the other hand, the reverse family of models is able to move past the noise. Third, our architecture implements what we term “double compression”: the combination of decreasing channel depth (512 → 256 → 128 → 64) with max pooling at each block creates simultaneous channel and spatial reduction. This is architecturally distinct from encoder-decoder networks such as SegNet (), which also reduce channel depth in their decoder pathway but do so while upsampling spatial resolution for segmentation tasks. In contrast, our architecture continues to downsample spatially while reducing channels, creating an aggressive information bottleneck optimized for classification rather than spatial reconstruction. This double compression forces the network to retain only the most salient biological features, such as gonad and digestive gland morphology, while discarding both stochastic speckle noise and fine-grained spatial details that do not contribute to sex discrimination. Fourth, we placed dropout at the end of the convolutional block, immediately before flattening. This strategic placement, combined with the high initial channel count (512 filters), provides regularization at the critical transition between spatial feature extraction and dense classification. The dropout layer acts as an additional noise filter, randomly zeroing activations that may have encoded speckle-specific patterns during training, thereby encouraging the network to learn robust representations of anatomically meaningful structures.

However, the ResNet model family still remained competitive with the reverse models. We propose that both models were able to combat overfitting, but did so in different ways. The ResNet poses several advantages over the VGG models, namely through their key skip connections, improving gradient flow and stabilizing training. The ResNet models enable and encourage identity layers, so that noisy layers can be nullified and skipped over via the skip connections, and is in general more suitable to a robust variety of applications. Thus the ResNet models also do not need a large dataset to efficiently generalize, and are generally able to combat noisy environments seen in ultrasound data.

Though we did not explore a combination of the reverse layers and the ResNet family (adding skip connections to the reverse CNN models), this represents a promising direction. Recent work by (Liu et al. 2025) demonstrated that their RevNet architecture, which combines the reverse filter topology with ResNet's residual block framework, achieved competitive or superior performance on standard benchmarks. A hybrid architecture combining our inverted channel progression (512 → 64) with skip connections could potentially leverage both the noise-filtering properties of progressive channel compression and the gradient flow benefits of residual connections, representing a logical next step for ultrasound classification tasks.

Our final and best performing models come from the YOLO family. We trained the YOLOv8x-cls and YOLOv11x-cls models from scratch, meaning we randomly initialized the model parameters/weights, instead of loading pretrained weights. At the time of this writing, the YOLOv8 models outperformed the YOLOv11 models, but we expect this may change as updates continue to be made to YOLOv11.

4.3 t-distributed Stochastic Neighbor Embedding (t-SNE) analysis

We first discuss the qualitative outputs of t-SNE on our dataset. Surprisingly, the most separated set of data belonged to the testing set, despite it having the lowest accuracy. There is significant overlap in the training and validation t-SNE plots, which is particularly surprising for such high accuracy results observed in our final YOLOv8 model. This suggests that the decision boundary is complex and high-dimensional, and has a very fine line for distinguishability. The model successfully separates classes in the original high-dimensional space, even though t-SNE fails to visualize this separation in 2D. This suggests that the learned features are simply still extremely complex. To further confirm this result, we found that the train, validation, and test silhouette scores were 0.0964, 0.0963, and 0.0795, respectively, on the intermediate feature values.

First, the intermediate features do not form well-separated clusters. This makes it particularly difficult for t-SNE to produce a well-separated graph, and it likely means that the decision boundaries are extremely fine-grained and subtle. Second, it aligns with human intuition: even trained experts find this classification task difficult without explicit visual cues. The model must rely on subtle and distributed visual signals, rather than clearly defined or localized features. Third, this task results in signals being distributed across features. In particular, the final features contain a variety of features which all have strong signals, as opposed to few neurons containing almost all of the discriminating features. This suggests that a variety of properties contribute together toward determining the output; possibly features such as gonad shape, brightness, and texture cues which provide only subtle hints and must work together.

4.4 Practical applications

In vivo sex determination in abalone is typically conducted via visual inspection, which can be unreliable due to gonad maturation states, yet our image classification models also identify structural features in the ultrasound images that are not apparent to the human eye. The activation of areas adjacent to the gonad raises the question of whether the proportion between the gonad and the digestive gland area is also important for sex determination. This capacity to highlight anatomically relevant structures—potentially as a consequence of the aggressive information bottleneck created by simultaneous channel and spatial reduction—suggests that our architecture may be particularly well-suited to soft tissue classification, given more optimization, in ultrasound where texture and regional morphology carry discriminative information. Although some of the models we trained with are over a decade old, they still perform well and above random chance when updated to meet current standards via more recent data augmentation, image preprocessing and regularization methods.

By coupling ultrasonography with machine learning-based inference on edge computing devices, this framework lays the groundwork for real-time, operator-independent sexing tools that can enhance spawning efficiency, reduce handling stress, and support scaling of recovery and sustainable propagation efforts for endangered abalone. For abalone programs where the number of available individuals is large, the ability to determine sex with accuracies above 80%, as observed for large red abalone after filtering low quality images, represents a substantial labor savings compared with manual evaluation of every animal, particularly when hundreds of individuals can be screened per hour with manual checks reserved for only a subset of animals or uncertain classifications. In practice, operators acquire multiple ultrasound frames per animal under the standardized tank conditions used in hatchery and aquaculture facilities; the envisioned deployment workflow would automatically screen frames for quality, classify sex on passing frames, and flag low-confidence predictions for manual review, reducing dependence on any single image being artifact-free. For captive breeding programs for threatened and endangered species, where broodstock numbers are limited and many individuals are immature, rapid non-lethal screening can be used to identify the relatively few animals that have reached reproductive maturity and prioritize them for spawning.

As future models are extended from sex classification to also approximate gonad maturity, we expect additional gains in spawning advancement, larval recruitment, and juvenile survival. Studies using ultrasonography in abalone have shown that individuals with higher gonad indices or greater gonad relative average thickness exhibit higher total and relative fecundity, higher fertilization and hatching rates, lower larval abnormality, and higher larval attachment than low index individuals, directly linking gonad development to reproductive performance and larval quality (Zou et al., 2025). Together with work demonstrating that ultrasound-derived gonad indices can be used to track maturation and time spawning events in red abalone and endangered California abalone species, these findings suggest that incorporating maturity prediction into machine learning models could further improve broodstock selection. In particular, preferentially selecting females with larger and more developed ovaries is likely to select for more mature ova, important for non-feeding, free-swimming larvae that depend entirely on maternal nutrient reserves. Improving our ability to identify and use such females should translate into improved larvae quality and recruitment outcomes in both conservation and production hatchery programs.

4.5 Limitations

Several limitations warrant consideration. First, the dataset consisted of only 34 large size class individuals, limiting statistical power and potentially inflating performance estimates. Second, models trained directly on small abalone performed poorly, and although transfer learning from large-animal weights improved test accuracy from approximately 55% to 60% (Supplementary Table S1), this remained insufficient for reliable classification. This suggests size-dependent morphological variation that was not captured, or perhaps their size limited the quality of the ultrasound on smaller animals leading to a larger amount of reflection and artifacts. Third, image quality significantly impacted classification accuracy; a substantial proportion of images were excluded due to artifacts, and we acknowledge that our current models were developed and evaluated using quality-filtered images from large individuals. Performance under variable imaging conditions and with smaller animals remains to be established. However, the intended deployment environments for this technology are conservation hatcheries and commercial aquaculture facilities, where abalone are imaged while submerged in tanks under standardized conditions, not in unpredictable field settings. The imaging protocol used in this study was performed at the breeding facility under the same conditions that would apply in routine hatchery operations. In practice, operators acquire multiple frames per animal, and we envision a deployment pipeline in which automated quality screening precedes sex classification (Section 4.6), with uncertain classifications flagged for manual review. This multi-frame, quality-aware workflow would substantially reduce dependence on individual image quality. Fourth, the study focused on a single species; transferability to endangered species (H. cracherodii, H. sorenseni) requires validation and significant time and resources. Fifth, the binary classification task does not address reproductive maturity assessment, which remains an important goal for spawning management and an interesting and natural follow up to this study.

4.6 Future directions

We identify five areas for future development. The first is expansion to additional species, particularly endangered abalone where this technology would have immediate conservation value. The second is integration of reproductive maturity assessment, potentially through multitask learning that jointly predicts sex and gonad development stages. The third is development of automated image quality control. Manual filtering in this study was performed by an operator with no prior ultrasonography training, demonstrating that the criteria are accessible, but the process remains subjective and would benefit from standardization. Deep learning-based quality assessment has been demonstrated for cardiac ultrasound () and high-frequency dermatological ultrasound (Kucharski et al., 2022), and a similar approach could automate the identification of artifact-laden or incomplete abalone ultrasound frames prior to sex classification. The computational cost of such a quality filter is minimal and compatible with the edge inference hardware described in Section 3.4, enabling a two-stage pipeline in which quality screening and sex classification run sequentially on a single device. The fourth is collection of larger datasets to enable more robust generalization across size classes and imaging conditions. The fifth is architectural exploration combining our inverted channel topology with residual connections; recent work by (Liu et al. 2025) demonstrated that their RevNet (not ResNet) architecture, which applies this decreasing filter pattern (512 → 256 → 128 → 64) within ResNet's residual framework, achieves competitive performance on standard benchmarks, suggesting that a hybrid approach could leverage both the noise-filtering properties of progressive compression and the gradient flow benefits of skip connections for ultrasound classification tasks.

Ultimately, our results demonstrate that computer vision models can classify red abalone sex from ultrasound images with useful accuracy, even when trained on a limited dataset. Although additional data are needed to improve robustness across size classes and image quality, the present models already achieve inference speeds that are compatible with on site use on embedded EdgeAI devices such as the Jetson Orin Nano, allowing biologists to perform multiple scans and classify animals in real time. As larger and more diverse ultrasound datasets are collected and practitioners become more familiar with AI embedded systems for classification, we expect both accuracy and reliability to improve, supporting more efficient broodstock management and spawning decisions in conservation and production aquaculture and increasing the throughput of endangered species rescue and breeding efforts.

5 Conclusions

We demonstrate that convolutional neural networks can classify red abalone sex from ultrasound images with 85.7% accuracy on held-out test data, establishing the feasibility of automated, non-invasive sex determination for this species. The YOLOv8 architecture outperformed classical CNN models (VGG, ResNet) and custom architectures, likely due to its efficient handling of limited training data. Notably, our finding that a reverse channel architecture (512 → 256 → 128 → 64) outperforms standard VGG on ultrasound data suggests that progressive channel compression may help CNNs overcome speckle noise, a hypothesis consistent with the Information Bottleneck principle (Tishby and Zaslavsky, 2015) and recently corroborated by (Liu et al. 2025), who demonstrated competitive performance with decreasing filter topologies on standard benchmarks. Our architecture implements what we term “double compression”: simultaneous channel reduction and spatial downsampling via max pooling, which is architecturally distinct from encoder-decoder networks like SegNet that reduce channels while upsampling spatially. This double compression, combined with strategic dropout placement at the end of the convolutional block, appears to enable the network to attend to biologically meaningful features such as gonad and digestive gland morphology rather than speckle artifacts, as evidenced by our feature activation maps. To our knowledge, this represents the first application of an inverted channel architecture with double compression to marine invertebrate tissue classification from ultrasound imaging, and the first demonstration of automated sex determination in abalone from ultrasound. By coupling ultrasonography with edge-deployable inference, this framework enables real-time, operator-independent sexing tools for both conservation breeding programs and commercial aquaculture. Future work should expand to endangered species, integrate reproductive maturity assessment, and explore hybrid architectures combining our inverted topology with residual connections.

Statements

Data availability statement

Source code is available in the following GitHub repositories, https://github.com/ESB-AI-Lab/AbaloneClassification, https://github.com/ESB-AI-Lab/red-abalone-sex-classification. The first link contains finalized code for the final training done for this manuscript, and the second link contains code for initial exploratory data analysis and training prior to finalizing a standard pipeline to be used across all model architectures. The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.

Ethics statement

The manuscript presents research on animals that do not require ethical approval for their study.

Author contributions

EAS: Formal analysis, Visualization, Data curation, Validation, Software, Investigation, Project administration, Writing – review & editing, Conceptualization, Writing – original draft, Methodology, Supervision. AT: Investigation, Writing – review & editing, Data curation, Software, Methodology, Visualization, Validation, Writing – original draft, Formal analysis. SY: Data curation, Software, Writing – original draft, Investigation, Visualization, Formal analysis. WG-N: Writing – original draft, Software, Methodology. KJ: Validation, Methodology, Writing – original draft, Software. GF: Writing – original draft, Data curation, Software. SL: Visualization, Formal analysis, Writing – original draft. YZ: Data curation, Writing – original draft. MM: Writing – original draft, Data curation. SEB: Visualization, Formal analysis, Validation, Conceptualization, Methodology, Writing – review & editing, Data curation, Writing – original draft, Investigation. AEF: Writing – original draft, Writing – review & editing, Investigation. JAG: Conceptualization, Validation, Project administration, Investigation, Supervision, Methodology, Writing – original draft, Resources, Funding acquisition, Writing – review & editing.

Funding

The author(s) declared that financial support was received for this work and/or its publication. Funding for a portion of this research was provided by the United States Navy, Commander Pacific Fleet N62473-19-0023 and N62473-22-2-0007. Additional funding was also provided by the National Oceanic and Atmospheric Administration Section 6 Grant NA19NMF4720103 to the California Department of Fish and Wildlife through subcontract No. P197003 to the University of California, Davis. Funding for black abalone rescue work was provided by the National Marine Sanctuary Foundation, the Bureau of Ocean Energy Management, National Oceanic and Atmospheric Administration, the National Marine Fisheries Service, and the California Ocean Protection Council.

Acknowledgments

We would like to thank Naomi Weimken for assistance in ultrasound imaging of abalone gonad state. This work used Expanse at San Diego Supercomputer Center, and Bridges-2 at Pittsburgh Supercomputing Center through allocation MCB180035 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support (ACCESS) program, which is supported by U.S. National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296 (). This work was supported (in part) by the University of California President's Postdoctoral Fellowship awarded to Solares.

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declared that generative AI was not used in the creation of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Supplementary material

The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/frai.2026.1794183/full#supplementary-material

References

  • 1

    AbdiA. H.LuongC.TsangT.AllanG.NouranianS.JueJ.et al. (2017). Automatic quality assessment of echocardiograms using convolutional neural networks: feasibility on the apical four-chamber view. IEEE Trans. Med. Imag. 36, 12211230. doi: 10.1109/TMI.2017.2690836

  • 2

    AltstattJ. M.AmbroseR. F.EngleJ. M.HaackerP. L.LafertyK. D.RaimondiP. T. (1996). Recent declines of black abalone Haliotis cracherodii on the mainland coast of central california. Marine Ecol. Prog. Series142, 185192. doi: 10.3354/meps142185

  • 3

    BadrinarayananV.KendallA.CipollaR. (2017). Segnet: a deep convolutional encoder-decoder architecture for image segmentation. IEEE Trans. Pattern Analy. Mach. Intellig. 39, 24812495. doi: 10.1109/TPAMI.2016.2644615

  • 4

    Barrera-HernandezR.Barrera-SotoV.Martinez-RodriguezJ. L.Rios-AlvaradoA. B.Ortiz-RodriguezF. (2023). “Towards abalone differentiation through machine learning,” in Applied Machine Learning and Data Analytics, eds. M. A. Jabbar, F. Ortiz-Rodríguez, S. Tiwari, and P. Siarry (Cham: Springer Nature Switzerland), 108118.

  • 5

    BarulinN. V. (2019). Using machine learning algorithms to analyse the scute structure and sex identification of sterlet acipenser ruthenus (acipenseridae). Aquacult. Res. 50, 28102825. doi: 10.1111/are.14233

  • 6

    BoernerT. J.DeemsS.FurlaniT. R.KnuthS. L.TownsJ. (2023). “Access: Advancing innovation: Nsf's advanced cyberinfrastructure coordination ecosystem: Services & support,” in Practice and Experience in Advanced Research Computing 2023: Computing for the Common Good, PEARC '23 (New York, NY: Association for Computing Machinery), 173176.

  • 7

    BolesS. E.NeylanI. P.Rogers-BennettL.GrossJ. A. (2022). Evaluation of gonad reproductive condition using non-invasive ultrasonography in red abalone (haliotis rufescens). Front. Marine Sci. 9:784481. doi: 10.3389/fmars.2022.784481

  • 8

    BolesS. E.Rogers-BennettL.BraggW. K.Bredvik-CurranJ.GrahamS.GrossJ. A. (2023). Determination of gonad reproductive state using non-lethal ultrasonography in endangered black (haliotis cracherodii) and white abalone (h. sorenseni). Front. Marine Sci. 10:1134844. doi: 10.3389/fmars.2023.1134844

  • 9

    BradskiG. (2000). The OpenCV library. Boulder, CO: Dr. Dobb's Journal of Software Tools.

  • 10

    ClarkA. (2015). Pillow (PIL fork) Documentation.

  • 11

    DabiriR. (2025). Non-invasive abalone sex classification from external measurements using interpretable machine learning. Int. J. Comp. Appl. 187, 6572. doi: 10.5120/ijca2025925985

  • 12

    EdieS. M.CollinsK. S.JablonskiD. (2023). High-throughput micro-ct scanning and deep learning segmentation workflow for analyses of shelly invertebrates and their fossils: Examples from marine bivalvia. Front. Ecol. Evol. 11:1127756. doi: 10.3389/fevo.2023.1127756

  • 13

    EstevaA.KuprelB.NovoaR. A.KoJ.SwetterS. M.BlauH. M.et al. (2017). Dermatologist-level classification of skin cancer with deep neural networks. Nature542, 115118. doi: 10.1038/nature21056

  • 14

    GardnerG. R.HarshbargerJ. C.LakeJ. L.SawyerT. K.PriceK. L.StephensonM. D.et al. (1995). Association of prokaryotes with symptomatic appearance of withering syndrome in black abalone Haliotis cracherodii. J. Invertebrate Pathol. 66, 111120. doi: 10.1006/jipa.1995.1072

  • 15

    GrahamC. A.ShamkhalichenarH.BrowningV. E.ByrdV. J.LiuY.Gutierrez-WingM. T.et al. (2022). A practical evaluation of machine learning for classification of ultrasound images of ovarian development in channel catfish (Ictalurus punctatus). Aquaculture552:738039. doi: 10.1016/j.aquaculture.2022.738039

  • 16

    GulshanV.PengL.CoramM.StumpeM. C.WuD.NarayanaswamyA.et al. (2016). Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs. JAMA316, 24022410. doi: 10.1001/jama.2016.17216

  • 17

    HahnK. (1989). Handbook of Culture of Abalone and Other Marine Gastropods. Boca Raton, FL: CRC Press.

  • 18

    HarrisC. R.MillmanK. J.van der WaltS. J.GommersR.VirtanenP.CournapeauD.et al. (2020). Array programming with NumPy. Nature585, 357362. doi: 10.1038/s41586-020-2649-2

  • 19

    HeK.ZhangX.RenS.SunJ. (2016). “Deep residual learning for image recognition,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (Las Vegas, NV: IEEE), 770778.

  • 20

    HobdayA. J.TegnerM. J. (2000). Status Review of White abalone (Haliotis sorenseni) Throughout Its Range in California and Mexico (NOAA Technical Memorandum NMFS-SWR-035). U.S. Department of Commerce; National Oceanic and Atmospheric Administration; National Marine Fisheries Service.

  • 21

    HobdayA. J.TegnerM. J.HaakerP. L. (2000). Over-exploitation of a broadcast spawning marine invertebrate: decline of the white abalone. Rev. Fish Biol. Fisher. 10, 493514. doi: 10.1023/A:1012274101311

  • 22

    HunterJ. D. (2007). Matplotlib: A 2D graphics environment. Comp. Sci. Eng. 9, 9095. doi: 10.1109/MCSE.2007.55

  • 23

    Joblib Development Team (2024). Joblib: Running Python Functions as Pipeline Jobs.

  • 24

    JocherG.ChaurasiaA.QiuJ. (2023). Ultralytics YOLOv8 (version 8.0.0) [Computer Software]. https://github.com/ultralytics/ultralytics

  • 25

    JocherG.QiuJ. (2024). Ultralytics YOLO11 (version 11.0.0) [Computer Software]. https://github.com/ultralytics/ultralytics

  • 26

    KarpovK.HaakerP.TanigichiI.Rogers-BennettL. (2000). Serial depletion and the collapse of the California abalone (Haliotis spp.) fishery. Can. Special Publicat. Fisher. Aquatic Sci. 130, 1124.

  • 27

    KucharskiD.KleczekP.Jaworek-KorjakowskaJ.GutowskiN.SzabóM. (2022). High-frequency ultrasound dataset for deep learning-based image quality assessment. Sensors22:1478. doi: 10.3390/s22041478

  • 28

    LaçiH.SevraniK.IqbalS. (2025). Deep learning approaches for classification tasks in medical X-ray, MRI, and ultrasound images: a scoping review. BMC Med. Imag. 25:156. doi: 10.1186/s12880-025-01701-5

  • 29

    LawleyA.HampsonR.WorrallK.DobieG. (2024). Analysis of neural networks for routine classification of sixteen ultrasound upper abdominal cross sections. Abdom. Radiol. 49, 651661. doi: 10.1007/s00261-023-04147-x

  • 30

    LiuH.BrailsfordT.BullL. (2025). Exploring filter placement in convolutional layer topologies based on ResNet for image classification. Mach. Vision Appl. 36:54. doi: 10.1007/s00138-025-01674-z

  • 31

    MaatenL.HintonG. (2008). Visualizing Data using t-SNE. J. Mach. Learn. Res. 9, 25792605.

  • 32

    NashW.SellersT.TalbotS.CawthornA.FordW. (1994). Abalone. Irvine: UCI Machine Learning Repository.

  • 33

    NeylanI. P.SwezeyD. S.BolesS. E.GrossJ. A.SihA.StachowiczJ. J. (2024). Within- and transgenerational stress legacy effects of ocean acidification on red abalone (Haliotis rufescens) growth and survival. Global Change Biol. 30:e17048. doi: 10.1111/gcb.17048

  • 34

    NMF (2020). Final Endangered Species Act Recovery Plan for Black Abalone (Haliotis cracherodii). National Marine Fisheries.

  • 35

    PaszkeA.GrossS.MassaF.LererA.BradburyJ.ChananG.et al. (2019). Pytorch: An imperative style, high-performance deep learning library. arXiv [preprint] arXiv:1912.01703. doi: 10.48550/arXiv.1912.01703

  • 36

    PedregosaF.VaroquauxG.GramfortA.MichelV.ThirionB.GriselO.et al. (2011a). Scikit-learn: machine learning in Python. J. Mach. Learn. Res. 12, 28252830.

  • 37

    PedregosaF.VaroquauxG.GramfortA.MichelV.ThirionB.GriselO.et al. (2011b). Scikit-learn: machine learning in Python. J. Mach. Learn. Res. 12, 28252830.

  • 38

    PetersH.RalphG. M.Rogers-BennettL. (2024). Abalones at risk: A global Red List assessment of Haliotis in a changing climate. PLoS ONE, 19:e0309384. doi: 10.1371/journal.pone.0309384

  • 39

    PetersH.Rogers-BennettL. (2018). IUCN Red List of Threatened Species: Haliotis sorenseni.

  • 40

    PetersH.Rogers-BennettL. (2020). IUCN Red List of Threatened Species: Haliotis walallensis.

  • 41

    PetersH.Rogers-BennettL. (2021a). IUCN Red List of Threatened Species: Haliotis corrugata.

  • 42

    PetersH.Rogers-BennettL. (2021b). IUCN Red List of ThreatenedSpecies: Haliotis cracherodii.

  • 43

    PetersH.Rogers-BennettL. (2021c). IUCN Red List of ThreatenedSpecies: Haliotis fulgens.

  • 44

    PetersH.Rogers-BennettL.De ShieldsR. M. (2021). IUCN Red List of Threatened Species: Haliotis walallensis.

  • 45

    Python Software Foundation (2023). Python Language Reference, version 3.10.x.

  • 46

    Rogers-BennettL.AquilinoK. M.CattonC. A.KawanaS. K.WalkerB. J.AshlockL. W.et al. (2016). Implementing a restoration program for the endangered white abalone (Haliotis sorenseni) in California. J. Shellfish Res. 35, 611618. doi: 10.2983/035.035.0306

  • 47

    Rogers-BennettL.DondanvilleR. F.KashiwadaJ. V. (2004). Size specific fecundity of red abalone (Haliotis rufescens): evidence for reproductive senescence J. Shellfish Res. 23, 553560.

  • 48

    Rogers-BennettL.DondanvilleR. F.MooreJ. D.Ignacio VilchisL. (2010). Response of red abalone reproduction to warm water, starvation, and disease stressors: implications of ocean warming. J. Shellfish Res. 29, 599611. doi: 10.2983/035.029.0308

  • 49

    Rogers-BennettL.KlamtR.CattonC. A. (2021). Survivors of climate driven abalone mass mortality exhibit declines in health and reproduction following kelp forest collapse. Front. Marine Sci. 8:725134. doi: 10.3389/fmars.2021.725134

  • 50

    Rogers-BennettL.RogersD.BennettW.EbertT. (2003). Modeling red sea urchin growth using six growth functions. Fishery Bullet101, 614626.

  • 51

    SimonyanK.ZissermanA. (2015). Very deep convolutional networks for large-scale image recognition. arXiv [preprint] arXiv:1409.1556. doi: 10.48550/arXiv.1409.1556

  • 52

    TishbyN.ZaslavskyN. (2015). “Deep learning and the information bottleneck principle,” in 2015 IEEE Information Theory Workshop (ITW), 1-5.

  • 53

    ZouW.HongJ.MaY.YuW.LiuY.AiC.et al. (2025). Application of ultrasonography in abalone gonadal evaluation. Aquacult. Fisher. 10, 10401049. doi: 10.1016/j.aaf.2024.04.001

Summary

Keywords

computer vision, conservation and production aquaculture, endangered species, high-throughput, machine learning, sustainable food production

Citation

Solares EA, Tong A, Yoo S, Gosline-Niheu W, Jacob K, Feliz G, Loecher S, Zhai Y, Mah M, Boles SE, Fagbohun AE and Gross JA (2026) Deep learning classification of reproductive tissue from ultrasound: sex determination in red abalone (Haliotis rufescens). Front. Artif. Intell. 9:1794183. doi: 10.3389/frai.2026.1794183

Received

23 January 2026

Revised

04 April 2026

Accepted

13 April 2026

Published

18 May 2026

Volume

9 - 2026

Edited by

Hanqi Zhuang, Florida Atlantic University, United States

Reviewed by

Dawei Sun, Zhejiang Academy of Agricultural Sciences, China

Patrick McGrath, Georgia Institute of Technology, United States

Updates

Copyright

*Correspondence: Edwin A. Solares,

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics