Abstract
The ability to perceive visual objects with various types of transformations, such as rotation, translation, and scaling, is crucial for consistent object recognition. In machine learning, invariant object detection for a network is often implemented by augmentation with a massive number of training images, but the mechanism of invariant object detection in biological brains—how invariance arises initially and whether it requires visual experience—remains elusive. Here, using a model neural network of the hierarchical visual pathway of the brain, we show that invariance of object detection can emerge spontaneously in the complete absence of learning. First, we found that units selective to a particular object class arise in randomly initialized networks even before visual training. Intriguingly, these units show robust tuning to images of each object class under a wide range of image transformation types, such as viewpoint rotation. We confirmed that this “innate” invariance of object selectivity enables untrained networks to perform an object-detection task robustly, even with images that have been significantly modulated. Our computational model predicts that invariant object tuning originates from combinations of non-invariant units via random feedforward projections, and we confirmed that the predicted profile of feedforward projections is observed in untrained networks. Our results suggest that invariance of object detection is an innate characteristic that can emerge spontaneously in random feedforward networks.
Introduction
Visual object recognition is a crucial function for animal survival. Human and primates can detect objects robustly, despite huge variations in the position, size, and viewing angles (; ; ; ; ; ). This challenging ability is thought to be based on invariant neural tuning in the brain—Neurons that selectively respond to various objects have been observed in higher visual areas, and these neurons showed invariant object representation across various types of transformation (; ; ; ; ; ; ; ; ). Behavioral- and neural-level observations of this function have led many researchers to raise the important question of how this invariance of object detection emerges.
Often, this invariant neural tuning has been considered to develop from the learning of various types of visual transformations (; ). With the notion that visual experience of natural objects contains numerous variants that transform depending on the viewing conditions, it has been suggested that the capability to detect objects invariantly can develop gradually when observers repeatedly see objects with a wide range of variations (). Notably, in the machine learning field, invariant object recognition is also implemented via learning with a massive dataset. In this case, data augmentation, a specialized method to increase the dataset volume, is often applied (; ; ) to generate images through linear transformations such as rotation, positional shifting, and flipping, as in natural visual experience. Then, invariant object recognition is achieved from the training of the augmented dataset with computer vision models (; ; ). In contrast to the above scenario, observations in newborn animals suggest the possibility of its emergence without learning: Human infants show a preference to faces despite variations of the size and rotations in depth (; ; ). In addition, newborn chicks can detect virtual objects from novel viewpoints (). These findings imply that invariant object detection arises without visual experience, but the developmental mechanism of this invariance in biological brains—how object invariance arises innately in the complete absence of learning—remains elusive.
A model study using a biologically inspired deep neural network (DNN) (; ) has been suggested as an effective approach to this problem (; ; ; ; , ; ; , ; ; ). DNNs, which consist of a stack of feedforward projections inspired by the hierarchical structure of the visual pathway, can be used as simplified model to investigate various visual functions. For instance, it was reported that a DNN trained to natural images can predict neural responses in the primate visual pathways from an early visual area (e.g., primary visual cortex, V1) to a higher visual area (e.g., inferior temporal cortex, IT) (; ). Recent studies also provided insight into the origin of functional tuning in the brain, by showing that units that selectively respond to numerosity, faces, and various types of objects among visual stimuli can arise in a randomly initialized DNN without any learning (; ).
By adopting a similar approach, here we show that object invariance can arise in completely untrained neural networks. Using AlexNet (), a model designed along the structure of the visual stream, we found that units selective to various visual objects were observed in a randomly initialized DNN and that these units maintained selectivity across a wide range of variations, such as the viewpoint, even without any visual training. We observed that a certain proportion of the units show an invariant tuning to viewpoint, while other groups of units show tuning to a specific viewpoint. Preferred feature images obtained from the reverse-correlation method showed that each specific viewpoint unit encodes a shape from a particular view of an object, while invariant viewpoint units encode inclusive features from specific units with different preferred angles. We found that invariant units emerge by homogenous projections from specific units in the previous layer in a random feedforward network. Finally, we confirmed that this innate invariance enables the network to perform an object-detection task under an enormous range of variations of viewpoints. Overall, our results suggest that invariant object detection can emerge spontaneously from the random wiring of hierarchal feedforward projections in an untrained DNN.
Results
Emergence of object selectivity in untrained networks
To investigate the emergence of invariant object selectivity in an untrained model network, we used AlexNet (), a biologically inspired DNN that models the structure of the ventral visual pathway. To find an object-selective response of an individual unit in the network, we investigated the responses of the final convolutional layer (Conv5), which is presumed to correspond to the IT domain of the brain. To simulate the condition of an untrained hierarchical network, we randomly initialized AlexNet using a standardized network initialization method (), by which the weights of the filters in every convolutional layer are randomly selected from a Gaussian distribution.
The stimulus set was designed to contain nine different object categories (e.g., Monitor, Bed, Chair, etc.) (Figure 1A). To define selective units for a specific target object, eight other class object sets and one scrambled set of the target object were used, following a previous experimental study () (see section “Materials and methods” for details). The images in each class were prepared by controlling the low-level features of the luminance, contrast, object location and object size (Supplementary Figure 1). Specifically, the pixel value distribution of the object image and background image were calibrated using the same Gaussian distribution (mean = 127.5, s.d. = 51.0), and the intra-class image similarity was also controlled at statistically comparable level. As this stimulus was given as input for the randomly initialized networks, the responses were measured in the Conv5 layer and an analysis of object selectivity was conducted (Figure 1B).
FIGURE 1
We found object-selective units that show higher responses to a specific class of target images (e.g., Toilet) than to other non-target class images and scrambled images (two-sided rank-sum test, P < 0.001) (Figure 1C). Among the nine object categories, we observed that object-selective units emerge mostly in a few objects categories (Figures 1D,E). We observed toilet-selective units (n = 565 ± 55 in 20 random networks, mean ± s.d.), sofa-selective units (n = 339 ± 68), and monitor-selective units (n = 294 ± 54) in the Conv5 layer (43,264 units; 13 × 13 × 256, Nx–position × Ny–position × Nchannel) (Figure 1D). In particular, the number of object-selective units was divided into large and small groups (Figure 1E, n = 20, two-sided rank-sum test, *P < 10–27). Large groups consisted of the toilet, sofa, and monitor groups (nunits = 400 ± 133) and small groups were the dresser, desk, bed, chair, nightstand, and table groups (nunits = 32 ± 32). Our previous study suggested that units selective to various visual objects can arise spontaneously from the simple configuration of the geometric components and that objects with a simple profile lead to a strong clustering of abstracted responses in the network, more likely to generate units selective to it (
We investigated the number and the selective index of object units across the convolutional layers and found that the number of object units increases when the convolutional layers become deeper (Supplementary Figure 3A). The object-selective index for a single unit also shows a strong tendency to increase across convolution layers, demonstrating that object tuning becomes sharper through the network hierarchy (Supplementary Figure 3B). Furthermore, we found that the responses of an untrained network measured in the deep layer (Conv5) were clustered as object classes in the latent space, while raw images do not cluster in the latent space (Figure 1F).
Invariance of object-selective units in untrained networks
Next, to investigate whether the observed object-selective units show viewpoint-invariant representations of an object image, we measured the responses of object-selective units to target objects and non-target objects with various viewpoint angles. To do this, a viewpoint-variant stimulus set was generated (Supplementary Figure 4) by rotating the viewpoint of 3D objects on the horizontal plane (Figure 2A). For each object, we rendered 13 variant images at different viewpoints between −90° and 90°. Then, we measured the responses of selective units to target objects and non-target objects with various viewpoint angles (Figure 2B). We found that units show selective responses when an object image within a certain threshold is presented (Figure 2C, left, n = 200, one-sided rank-sum test, Toilet at 0° vs. Non-toilet, *P < 10–13; Toilet at 45° vs. Non-toilet, **P < 10–5), while the units did not show selectivity when an object image at a larger viewpoint angle was given (Figure 2C, left, Toilet at 90° vs. Non-toilet, NS, P = 0.492). Hence, the selectivity of object-selective units is maintained within a limited effective range (Figure 2C, right).
FIGURE 2

Viewpoint-invariant object selectivity observed in an untrained network: (A) An object renderer generates object images at various viewpoints rotating in a horizontal orbit. (B) The viewpoint-varying object stimulus was generated within the viewpoint range of –90° to +90° with 13 steps. The responses of the object-selective unit were measured in the final convolutional layer of an untrained AlexNet. (C) Viewpoint-invariant responses of a selective unit for viewpoint-rotated object stimulus; the tuning of a single toilet-selective unit shows a wide range of viewpoint invariance. Shaded areas and error bars represent the standard error of 200 images (n = 200, one-sided rank-sum test, Toilet at 0° vs. Non-toilet, *P < 10−13; Toilet at 45° vs. Non-toilet, **P < 10−5; NS, P = 0.492). (D) The effective range of the average responses of object units and the pixel-wise correlation of a raw image (n = 200, one-sided rank-sum test, P < 0.05). (E) Average response of object-selective units for an object stimulus at different viewpoint rotations. The arrow indicates the effective range of the selective response, and the shaded area between dashed lines indicates the effective range of the raw-image correlation. Shaded areas represent the standard deviation of 200 images. (F) Comparison of effective ranges between the selective response and raw-image correlation in each object-selective unit (n = 20, two-sided rank-sum test; Toilet unit, *P < 10−4; Sofa unit, *P < 10−4; Monitor unit, *P < 10−4). Error bars indicate the standard deviation of 20 random networks.
To investigate the effective range that maintains the selectivity of object units quantitatively, we investigated the responses of selective units with a viewpoint between −90° and 90° and estimated the boundary of the viewpoint variation around which target-object tuning is lost. For example, we observed that object tuning of toilet units was retained when the viewpoint change was within 105° (Figure 2D, left, n = 200, one-sided rank-sum test, P < 0.05). Then, to verify whether the viewpoint invariance of an object-selective unit simply arises due to the similarity of the object shape upon a change of the viewpoint, we estimated the pixel-wise raw-image correlations between object images from a front view and a rotated view. We compared the effective ranges of viewpoint invariance between the selective responses and the image correlations (Figure 2D, right). For toilet units, we observed that the effective range of the selective responses is significantly wider than that of the image correlation (Figure 2E, Toilet units). Similarly, this tendency was commonly observed in other object-selective units (Figures 2E,F, n = 20, two-sided rank-sum test; Toilet unit, *P < 10–4; Sofa unit, *P < 10–4; Monitor unit, *P < 10–4). This result suggests that the observed invariance is not simply due to the similarity of the object images at different viewpoints but is a characteristic of object-selective units in untrained networks. To find the origin of the invariance in an untrained network, we also examined the single-unit-level characteristics of invariance. We found that each unit shows considerable variations in the response characteristics when a target object image with various viewpoints is given as the input. In particular, each unit shows various effective ranges (Figures 3A,B, left). Considering the definition of viewpoint invariance, we presumed that the top 30% of units were “viewpoint-invariant” units and the bottom 30% units were “viewpoint-specific” units in the subsequent analyses. Indeed, we observed that each tentative viewpoint-specific unit has various preferred angles; i.e., they only respond to a particular view of an object (Figure 3B, right).
FIGURE 3

Single-unit-level analysis of invariance: (A) Viewpoint tuning curves of two sample toilet units: a unit with a wide effective range (Top 30%) and a unit with a narrow effective range (Bottom 30%). (B) Histogram of the invariance effective range of each unit and histogram of the preferred angle of units with a narrow effective range (Bottom 30%). (C) Responses of individual viewpoint-specific and viewpoint-invariant toilet-selective units in the Conv5 layer (Viewpoint-specific units, one-way ANOVA with single peak filtering, P < 0.05; Viewpoint-invariant units, one-way ANOVA, P > 0.05). (D) Average tuning curves of viewpoint-specific units (n = 147) and viewpoint-invariant units (n = 95) in an untrained network. Shaded areas represent the standard error of each type of unit. (E) Overall process of the preferred feature image (PFI) of target units in Conv5 of untrained networks using a reverse-correlation analysis (
To verify our conjecture that “viewpoint-specific” and “viewpoint-invariant” units exist and can be classified according to the observed effective range of each unit (Figures 3A,B), we investigated the responses of object-selective units for object images with different viewpoints. Target object images with a viewpoint between −60° and 60° (five steps, 50 images per viewpoint class) were presented to the network, and the responses were measured. Indeed, we observed that there are units that only respond to a particular viewpoint image (Figure 3C, Viewpoint-specific, one-way ANOVA with single peak filtering, P < 0.05) and units that respond invariantly to any viewpoint image (Figure 3C, Viewpoint-invariant, one-way ANOVA, P > 0.05). Viewpoint-specific units show highly tuned responses to one preferred viewpoint angle, while viewpoint-invariant units show a flat tuning curve to any viewpoint (Figure 3D).
Next, to visualize the distinct tuning features of viewpoint-specific and viewpoint-invariant units, we used a reverse-correlation method (
The feedforward model can explain the spontaneous emergence of invariance
To validate the hypothesis that viewpoint-invariant units originate from the projection of viewpoint-specific units in the previous layer, we backtracked projections of the units from the source layer (Conv4) to the target layer (Conv5) and examined the weights of connected viewpoint-specific units. First, we confirmed that viewpoint-specific (n = 765 ± 102) and viewpoint-invariant toilet-selective units (n = 130 ± 28) exist in Conv4 as well as in Conv5 (Viewpoint-specific, n = 504 ± 81, Viewpoint-invariant, n = 96 ± 16). We confirmed that the viewpoint-specific units in Conv5 receive stronger input from units with the same object tuning than from other units in Conv4 (Figure 4A, left and middle, n = 20, two-sided rank-sum test, *P < 10–7). In more detail, the viewpoint-specific units in Conv5 receive inputs from Conv4 units strongly biased to a particular viewpoint angle (Figure 4A, right, n = 20, one-way ANOVA, *P < 10–11). This tendency of a strongly biased weight also appeared in other preferred viewpoints. The connectivity between viewpoint-specific units with the same preferred angle in the source and target layers showed significantly high weights compared to other projection directions (Figure 4B, n = 20, one-way ANOVA, *P < 0.05).
FIGURE 4

Emergence of viewpoint invariance based on unbiased projection from viewpoint-specific units: (A) Connectivity diagram (left); averaged weight from target object units and non-target object units (Conv4) to viewpoint-specific units (Conv5) (middle, n = 20, two-sided rank-sum test, *P < 10−7); averaged weight from viewpoint-specific units (Conv4) to viewpoint-specific units (Conv5) (right, n = 20, one-way ANOVA, *P < 10−11). Error bars indicate the standard error of 20 random networks. (B) Heatmap of weights between specific units from the source layer and specific units from the target layer. This heatmap shows biased input to specific units (n = 20, one-way ANOVA, *P < 0.05). (C) Connectivity diagram (left); averaged weight from target object units and non-target object units (Conv4) to viewpoint-invariant units (Conv5) (middle; n = 20, two-sided rank-sum test, *P < 10−7); and averaged weight from viewpoint-specific units (Conv4) to viewpoint-invariant units (Conv5) (right, n = 20, one-way ANOVA, NS, P = 0.313). Error bars indicate the standard error of 20 random networks. (D) The homogeneous index of the input projection weight that connected viewpoint-specific units (Conv4) to viewpoint-invariant units (Conv5) and viewpoint-specific units (Conv5), respectively (n = 20, two-sided rank-sum test, *P < 0.05). Error bars indicate the standard deviation of 20 random networks.
We also found that viewpoint-invariant units in Conv5 are strongly connected to units with the same object selectivity in Conv4 (Figure 4C, left and middle, n = 20, two-sided rank-sum test, *P < 10–7), as in the case of viewpoint-specific units. However, the viewpoint-invariant units in Conv5 receive homogeneous inputs from specific units in Conv4 units with various preferred viewpoint angles (Figure 4C, right, n = 20, one-way ANOVA, NS, P = 0.313). To estimate the degree of homogeneity in the projection weights, we defined the homogenous index as the inverse of the standard deviation of the average weight connected to specific units with different viewpoints in the source layer. The index of the average weight connected to viewpoint-invariant units is significantly higher than that of the average weight connected to the viewpoint-specific units, indicating an unbiased input to the viewpoint-invariant units (Figure 4D, n = 20, two-sided rank-sum test, Toilet; Invariant units vs. Specific units, *P < 0.05). This tendency was also observed in units with other object tunings (Figure 4D, n = 20, two-sided rank-sum test, Sofa and Monitor; Invariant units vs. Specific units, *P < 0.05; Supplementary Figure 6). This implies that observed viewpoint invariance of object tuning can originate from hierarchical random feedforward projections.
To verify this developmental model further, we revisited earlier observations of invariant object tuning in the monkey IT which reported that neurons in the higher layer in the hierarchy show increased invariance (from the ML to the AM area) (
From this result, we investigated whether this connectivity profile induces an increased trend of invariance across layers in our model neural network and found that such layer-specific characteristics of viewpoint invariance also emerge in the untrained network we used. We observed that the level of invariance increased along the network hierarchy (Supplementary Figure 7C). To quantify these invariance characteristics, we introduced an invariance index of units, defined as the inverse of the standard deviation of responses across different viewpoints. We observed an increase in the invariance index of selective units higher up in the hierarchy in the untrained AlexNet (Supplementary Figure 7D). The viewpoint-invariance index in Conv4 is significantly higher than that in Conv3 (n = 20, two-sided rank-sum test, *P < 10–7). Also, the viewpoint-invariance index in Conv5 is significantly higher than that in Conv4 (n = 20, two-sided rank-sum test, **P < 10–7). This increasing tendency of the viewpoint-invariance index along the network hierarchy is also observed in other object-selective units (n = 20, two-sided rank-sum test; Sofa, *P < 10–7, **P < 10–7; Monitor, *P < 10–7, **P < 10–7). In addition, we confirmed that the same increasing tendency of the number of invariant units along the network hierarchy exists across various object tunings (Supplementary Figure 7E, n = 20, two-sided rank-sum test; Toilet, *P < 0.05, **P < 0.001; Sofa, *P < 0.05; Monitor, *P < 10–5, **P < 10–6). These results suggest that our model provides a plausible scenario for understanding the spontaneous emergence of invariant object selectivity in untrained networks, which is supported by previous experimental observations of neural tunings.
Innate invariance enables invariant object detection without data-augmented learning
Next, we tested whether this innate invariance in untrained networks enables the network to perform the invariant object-detection task without learning. We expected that the information given by invariant object units is sufficient to detect an object while the viewpoint of the given object image varies, and in particular, that viewpoint-invariant units play a key role in enabling invariant object detection. To confirm this hypothesis, we designed two different methods to train an SVM which classifies whether or not a given image is a target object, using unit responses to stimulus given (Figure 5A). In the first case (Train 1), the SVM is trained using an object image with various viewpoints to train the SVM, while it is trained using the object image only with a center-fixed viewpoint in the second condition (Train 2). After training, object images with various viewpoints were used for the test session (Figure 5B, left). We performed this process using both invariant units and specific units.
FIGURE 5

Invariantly tuned unit responses enable invariant object detection: (A) Overall process of the object-detection task and SVM classifier using the responses of object-selective units. To train the SVM classifier, 30 images of a target object and 30 images of a non-target object were used. Among the 60 images, 40 images were randomly sampled for training and the remaining 20 were used to test the task performance. The responses of untrained network units for these images are obtained and used to train and test the SVM to classify whether or not the given image is the target object. (B) Train 1 uses object images with various viewpoints to train the SVM, while Train 2 uses object images only with a center-fixed viewpoint. For the test SVM, object images with various viewpoints are used (n = 20, two-sided rank-sum test, Invariant, NS, P = 0.735; Specific, *P < 0.001). (C) Performance of the Train 2 method using invariant units, specific units, and non-selective units (n = 20, two-sided rank-sum test, Invariant vs. Specific, *P < 10−7; Specific vs. Chance level, **P < 10−7; Non-selective vs. Chance level, NS, P = 0.116). The upper dashed line represents the performance when all units in Conv5 are used, and the lower dashed line indicates the chance level of the task. (D) The task in various test-image-variation-range conditions. The test images were randomly sampled within the given viewpoint variation range. The Train 2 performances of invariant units, all specific units, specific units only with 0° preferred, and non-selective units were assessed. (E) Comparison of performances across the types of object units. The performance for each case was measured using test images with a 180° variation range (n = 20, two-sided rank-sum test, Invariant vs. Specific with center view, *P < 0.05; All specific vs. Specific with center view, **P < 0.01; Specific with center view vs. Non-selective, ***P < 10−6). Shaded areas and error bars indicate the standard deviation of 20 random networks.
We found that the performances of the SVM using invariant units only and those of the SVM using specific units only are noticeably different (Figure 5B, right, Invariant, n = 20, two-sided rank-sum test, NS, P = 0.735; Specific, n = 20, two-sided rank-sum test, *P < 0.001). Invariant units show the same level of performance regardless of training with various viewpoints, implying that the information given by invariant units is sufficient to detect images with varying viewpoints. In contrast, specific units show significantly lower performance outcomes when trained only with a fixed viewpoint (Train 2). Hence, we investigated in more depth the performances of SVMs trained in the second condition (Train 2) using different groups of units (Figure 5C, n = 20, two-sided rank-sum test, Invariant vs. Specific, *P < 10–7; Specific vs. Chance level, **P < 10–7; Non-selective vs. Chance level, NS, P = 0.116). First, the SVM using the responses of invariant units shows significantly high performance compared to the SVM using the responses of specific units. Second, the SVM using invariant units shows the same level of performance compared to when it is trained with all units in the Conv5 layer, implying that the invariance of the untrained network mainly relies on invariant units. To confirm that invariance enables invariant object detection in a wide variation range, we tested this concept in different viewpoint-variation ranges (Figure 5D, left). When the viewpoint variation range in the test set becomes wider, the SVM using specific units rapidly loses its performance capabilities, whereas the SVM using invariant units maintains its high-performance outcomes (Figure 5D, right). Interestingly, the SVM using all specific units with different preferred angles outperforms the SVM using specific units only with the same preferred angle. This trend was also observed using other object-selective units (Figure 5E, n = 20, two-sided rank-sum test, Invariant vs. Specific with center view, *P < 0.05; All specific vs. Specific with center view, **P < 0.01; Specific with center view vs. Non-selective, ***P < 10–6). These results demonstrate that invariance in an untrained network enables an object-detection task with images with various viewpoints for a wide variation range, even without data-augmented learning.
Discussion
We showed that selectivity to various object emerges in randomly initialized networks and that this selectivity is robustly preserved even as the viewpoint changes significantly in the complete absence of learning. Furthermore, we found that the invariant tuning property can arise solely from the distribution of weights in feedforward projections. These results suggest that the statistical complexity of hierarchical neural network circuits allows the initial development of selectivity as well as invariance to various objects across a wide range of transformations.
Our results imply that innate invariance of object selectivity can arise from random feedforward projection, but this does not mean that there is no effect of experience on the development of this function. In fact, observations in various animals support the contention that this invariant function is affected by visual experience. In pigeons, the ability to detect objects across different variations in the viewing conditions is enhanced gradually during the visual training process (
Although the current study investigated only viewpoint invariance, we anticipate that invariance to other types of image transformations, such as position, size, and rotation, can also emerge spontaneously in untrained neural networks. Previous studies using an untrained DNN provide supporting evidence.
We proposed a method of generating invariance without learning, in contrast to previous approaches that implement the same function by relying on a massive training process. In the machine learning field, invariant object recognition has been implemented by learning a great many images. To learn invariant object features, the data-augmentation method is often applied (
In summary, we conclude that invariance of object selectivity can arise from the statistical variance of randomly wired bottom-up projections in untrained hierarchical neural networks. Our findings may provide new insight into the developmental mechanism of innate cognitive functions in biological and artificial neural networks.
Materials and methods
Untrained AlexNet
Currently, DNN models, which have a biologically inspired hierarchical structure, provide an effective approach for investigating functions in the brain (
Following earlier work, we used a randomly initialized (untrained) AlexNet (
Viewpoint-controllable object stimulus renderer
There are a few well-known objects image datasets, such as ImageNet (
ModelNet10 (
Stimulus dataset
We prepared three types of visual stimulus datasets specialized to each task. (1) Object dataset (Supplementary Figure 1A): This set was used to find units that selectively respond to a particular object class. It contains nine object classes (bed, chair, desk, dresser, nightstand, monitor, sofa, table, and toilet), and 200 images are prepared in each object class. To render the images of the object dataset, the viewpoint variation angle was randomly set between −30° and +30°. In the object dataset, brightness and contrast of the images are precisely controlled to be equal across object classes (Supplementary Figures 1B,C). In addition, the intra-class similarity of the images in each object category was calibrated at a statistically comparable level (Supplementary Figure 1D). (2) Viewpoint dataset (Supplementary Figure 4A): This set was used to test the viewpoint-invariant characteristics of the object-selective units. This dataset consists of 13 subsets which have different viewpoints from −180° to +180° on a linear scale. It contains 250 different object identities in an object class. Among them, 200 object identities are identical to those used in the object dataset. They were used to analyze the viewpoint-invariant characteristics of the object-selective units quantitatively. The remaining 50 object identities were used not to find object-selective units but to distinguish object-selective units with or without viewpoint invariance. In the viewpoint dataset, the luminance and contrast are also controlled (Supplementary Figures 4B,C). (3) SVM dataset: This set was used to train and test the SVM that performs the object-detection task. It contains 60 different object identities in an object class, which were not used for finding object-selective units. Specifically, it consists of 18 subsets with different viewpoint variations ranging from 0° to 180°. For example, a subset with a 180° viewpoint variation range contains images that show different viewpoints of objects within −90° and +90°.
Analysis of responses of the network units
Using the totally untrained AlexNet, we measured the responses of the target layer for each designed stimulus. For each response from the target convolution layers, each unit of an activation map was separately recorded for different classes of the stimulus. Based on our previous study, object-selective units were defined as units that showed a significantly greater mean response to target object images compared to those of non-target object images (P < 0.001, two-sided rank-sum test). To analyze the responses of each unit, it was necessary to regularize the raw response. To normalize the raw response, we used the z-scoring method. Furthermore, we used a trick in the z-score in that we subtracted from . indicates the response for an object class that leads to the second maximum response for that unit. Therefore, if the z-scored response is higher than zero, our unit shows a higher raw response to the target object than to the second maximum object, indicating selectivity.
To quantify the degree of tuning, an object selectivity index (OSI) of a single unit was defined using the follow formula. This index is modified from the face-selective index (FSI), which defined in previous experimental research (
is the average response to target-object images and is the average response to all non-target-object images. A higher OSI indicates fine tuning and an OSI of zero indicates equal responses to target and non-target object images.
Among the object-selective units, we defined a viewpoint-invariant unit as a unit for which the response was not significantly different (one-way ANOVA, P > 0.05) for all viewpoint classes. Similarly, viewpoint-specific units are defined as a unit for which the response was significantly high for one preferred viewpoint class (one-way ANOVA with single peak filtering, P < 0.05). For this, we detected a peak by thresholding the value of the average signal plus the standard deviation, as often done in the field of signal processing.
To measure the invariant index quantitatively, we calculated the inverse of the standard deviation of the average responses for images within each viewpoint class.
is the average response to a viewpoint class and μ is average response for all viewpoint classes. n is total number of viewpoint classes.
Preferred feature image analysis
To achieve the preferred input features of each target unit, we estimated the receptive field of units using the reverse correlation method (
Connectivity analysis
To investigate the connectivity between object-selective units across convolutional layers, we backtracked projections of the units from the source layer (Conv4) to the projection layer (Conv5). This backtracking process is opposite of the group convolution process. To backtrack the origin of a unit in the projection layer, we investigated all connected weights and units in the source layers.
To measure the degree of homogeneity in the input projection weight to a single target unit, the homogeneous index was defined as
where is the average weight from specific units in the source layer to a unit in the projection layer and μ is the average weight from all specific units. n is the total number of viewpoint-specific units with different preferred angles. To compare the unbiased properties of specific and invariant units, we normalized the homogenous index so that the average index value of viewpoint-specific units reaches unity.
Object-detection task
To validate viewpoint-invariant object-selectivity that spontaneously emerges in an untrained DNN, we trained a support vector machine (SVM) using the responses of object-selective units with two types of training. For Train 1, target-object (n = 40) or non-target-object (n = 40) images, which shows different viewpoints of objects within a range of –60° and +60° were randomly presented to the networks, and the observed responses of the Conv5 layer were used to train the SVM. For Train 2, most of the processes are nearly identical compared Train 1, but the only difference is in how the train images are presented. We prepared target-object and non-target object images without viewpoint variation (front-view only). After training the SVM, we investigated the performance with the responses of object-selective units for a stimulus with viewpoint variation. Here, target-object (n = 20) or non-target-object (n = 20) images were also randomly presented to the networks, and the responses from the Conv5 layer was used to test the SVM.
Statements
Data availability statement
The stimulus datasets and the MATLAB codes for this study are available at https://github.com/vsnnlab/Invariance.
Author contributions
S-BP conceived of the project. JC, SB, and S-BP designed the model and wrote the manuscript. JC performed the simulations. JC and SB analyzed the data. All authors contributed to the article and approved the submitted version.
Funding
This work was supported by a grant from the National Research Foundation of Korea (NRF) funded by the Korean government (MSIT) (Nos. NRF-2022R1A2C3008991, NRF-2021M3E5D2A01019544, and NRF-2019M3E5D2A01058328), the Singularity Professor Research Project of KAIST, and the KAIST Undergraduate Research Participation (URP) program (to S-BP).
Conflict of interest
The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
Supplementary material
The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fncom.2022.1030707/full#supplementary-material
References
1
AparicioP. L.IssaE. B.DiCarloJ. J. (2016). Neurophysiological organization of the middle face patch in macaque inferior temporal cortex.J. Neurosci.3612729–12745. 10.1523/JNEUROSCI.0237-16.2016
2
Apurva Ratan MurtyN.ArunS. P. (2015). Dynamics of 3D view invariance in monkey inferotemporal cortex.J. Neurophysiol.1132180–2194. 10.1152/jn.00810.2014
3
BaekS.ParkY.PaikS.-B. (2020). Sparse long-range connections in visual cortex for cost-efficient small-world networks.bioRxiv [Preprint]. 10.1101/2020.03.19.998468
4
BaekS.SongM.JangJ.KimG.PaikS.-B. (2021). Face detection in untrained deep neural networks.Nat. Commun.121–15. 10.1038/s41467-021-27606-9
5
BiedermanI. (1987). Recognition-by-components: A theory of human image understanding.Psychol. Rev.94:115. 10.1037/0033-295X.94.2.115
6
BoninV.HistedM. H.YurgensonS.Clay ReidR. (2011). Local diversity and fine-scale organization of receptive fields in mouse visual cortex.J. Neurosci.3118506–18521. 10.1523/JNEUROSCI.2974-11.2011
7
CadieuC. F.HongH.YaminsD. L. K.PintoN.ArdilaD.SolomonE. A.et al (2014). Deep neural networks rival the representation of primate it cortex for core visual object recognition.PLoS Comput. Biol.10:e1003963. 10.1371/journal.pcbi.1003963
8
ChenW.TianL.FanL.WangY. (2019). “Augmentation invariant training,” in Proceedings of the 2019 IEEE/CVF international conference on computer vision workshop (ICCVW), (Seoul: IEEE), 2963–2971. 10.1109/ICCVW.2019.00358
9
ConnorC. E.BrincatS. L.PasupathyA. (2007). Transformation of shape information in the ventral pathway.Curr. Opin. Neurobiol.17140–147. 10.1016/j.conb.2007.03.002
10
DiCarloJ. J.ZoccolanD.RustN. C. (2012). How does the brain solve visual object recognition?Neuron73415–434. 10.1016/j.neuron.2012.01.010
11
FöldiákP. (1991). Learning invariance from transformation sequences.Neural Comput.3194–200. 10.1162/neco.1991.3.2.194
12
FreiwaldW. A.TsaoD. Y. (2010). Functional compartmentalization and viewpoint generalization within the macaque face-processing system.Science330845–851. 10.1126/science.1194908
13
HungC. P.KreimanG.PoggioT.DiCarloJ. J. (2005). Fast readout of object identity from macaque inferior temporal cortex.Science310863–866. 10.1126/science.1117593
14
IchikawaH.NakatoE.IgarashiY.OkadaM.KanazawaS.YamaguchiM. K.et al (2019). A longitudinal study of infant view-invariant face processing during the first 3–8 months of life.Neuroimage186817–824. 10.1016/j.neuroimage.2018.11.031
15
ItoM.TamuraH.FujitaI.TanakaK. (1995). Size and position invariance of neuronal responses in monkey inferotemporal cortex.J. Neurophysiol.73218–226. 10.1152/jn.1995.73.1.218
16
JangJ.SongM.PaikS. B. (2020). Retino-cortical mapping ratio predicts columnar and salt-and-pepper organization in mammalian visual cortex.Cell Rep.303270.e–3279.e. 10.1016/j.celrep.2020.02.038
17
KaufmanL.RousseeuwP. J. (2009). Finding groups in data: An introduction to cluster analysis.Hoboken, NJ: John Wiley & Sons.
18
KimG.JangJ.BaekS.SongM.PaikS.-B. (2021). Visual number sense in untrained deep neural networks.Sci. Adv.7:eabd6127. 10.1126/sciadv.abd6127
19
KimJ.SongM.JangJ.PaikS. B. (2020). Spontaneous retinal waves can generate long-range horizontal connectivity in visual cortex.J. Neurosci.406584–6599. 10.1523/JNEUROSCI.0649-20.2020
20
KobayashiM.OtsukaY.KanazawaS.YamaguchiM. K.KakigiR. (2012). Size-invariant representation of face in infant brain: An fNIRS-adaptation study.Neuroreport23984–988. 10.1097/WNR.0b013e32835a4b86
21
KrizhevskyA.SutskeverI.HintonG. E. (2012). Imagenet classification with deep convolutional neural networks.Adv. Neural Inf. Process. Syst.251097–1105.
22
LeCunY.BottouL.OrrG.MullerK.-R. (1998). Ef?cient backprop. Neural networks tricks trade.New York, NY: Springer. 10.1007/3-540-49430-8_2
23
LiN.CoxD. D.ZoccolanD.DiCarloJ. J. (2009a). What response properties do individual neurons need to underlie position and clutter “invariant” object recognition?J. Neurophysiol.102360–376. 10.1152/jn.90745.2008
24
LiN.DiCarloJ. J. (2010). Unsupervised natural visual experience rapidly reshapes size-invariant object representation in inferior temporal cortex.Neuron671062–1075. 10.1016/j.neuron.2010.08.029
25
LiY.PizloZ.SteinmanR. M. (2009b). A computational model that recovers the 3D shape of an object from a single 2D retinal representation.Vision Res.49979–991. 10.1016/j.visres.2008.05.013
26
LogothetisN. K.PaulsJ.BülthoffH. H.PoggioT. (1994). View-dependent object recognition by monkeys.Curr. Biol.4401–414. 10.1016/S0960-9822(00)00089-0
27
O’GaraS.McGuinnessK. (2019). “Comparing data augmentation strategies for deep image classification,” in Proceedings of the irish machine vision & image processing conference, (Dublin: Technological University Dublin).
28
PaikS. B.RingachD. L. (2011). Retinal origin of orientation maps in visual cortex.Nat. Neurosci.14919–925. 10.1038/nn.2824
29
ParkY.BaekS.PaikS. B. (2021). A brain-inspired network architecture for cost-efficient object recognition in shallow hierarchical neural networks.Neural Netw.13476–85. 10.1016/j.neunet.2020.11.013
30
PerrettD. I.OramM. W.HarriesM. H.BevanR.HietanenJ. K.BensonP. J.et al (1991). Viewer-centred and object-centred coding of heads in the macaque temporal cortex.Exp. Brain Res.86159–173. 10.1007/BF00231050
31
PintoN.CoxD. D.DiCarloJ. J. (2008). Why is real-world visual object recognition hard?PLoS Comput. Biol.4:0151–0156. 10.1371/journal.pcbi.0040027
32
PoggioT.UllmanS. (2013). Vision: Are models of object recognition catching up with the brain?Ann. N.Y. Acad. Sci.130572–82. 10.1111/nyas.12148
33
Ratan MurtyN. A.ArunS. P. (2017). A balanced comparison of object invariances in monkey IT neurons.eNeuro41–10. 10.1523/ENEURO.0333-16.2017
34
RussakovskyO.DengJ.SuH.KrauseJ.SatheeshS.MaS.et al (2015). Imagenet large scale visual recognition challenge.Int. J. Comput. Vis.115211–252. 10.1007/s11263-015-0816-y
35
SailamulP.JangJ.PaikS. B. (2017). Synaptic convergence regulates synchronization-dependent spike transfer in feedforward neural networks.J. Comput. Neurosci.43189–202. 10.1007/s10827-017-0657-5
36
ShortenC.KhoshgoftaarT. M. (2019). A survey on image data augmentation for deep learning.J. Big Data61–48. 10.1186/s40537-019-0197-0
37
SimardP. Y.SteinkrausD.PlattJ. C. (2003). “Best practices for convolutional neural networks applied to visual document analysis,” in Proceedings of the international conference on document analysis and recognition, ICDAR 2003-Janua, (Edinburgh: IEEE), 958–963. 10.1109/ICDAR.2003.1227801
38
SimonyanK.ZissermanA. (2015). “Very deep convolutional networks for large-scale image recognition,” in Proceedings of the 3rd international conference learning representations ICLR 2015, San Diego - Conf.1–14.
39
SongM.JangJ.KimG.PaikS. B. (2021). Projection of orthogonal tiling from the retina to the visual cortex.Cell Rep.34:108581. 10.1016/j.celrep.2020.108581
40
StiglianiA.WeinerK. S.Grill-SpectorK. (2015). Temporal processing capacity in high-level visual cortex is domain specific.J. Neurosci.3512412–12424. 10.1523/JNEUROSCI.4822-14.2015
41
TanakaK. (1996). Inferotemporal cortex and object vision.Annu. Rev. Neurosci.19109–139. 10.1146/annurev.ne.19.030196.000545
42
TuratiC.BulfH.SimionF. (2008). Newborns’ face recognition over changes in viewpoint.Cognition1061300–1321. 10.1016/j.cognition.2007.06.005
43
van der MaatenL.HintonG. (2008). Visualizing data using t-SNE.J. Mach. Learn. Res.92579–2605.
44
WallisG.RollsE. T. (1997). Invariant face and object recognition in the visual system.Prog. Neurobiol.51167–194. 10.1016/S0301-0082(96)00054-8
45
WatanabeS. (1999). Enhancement of viewpoint invariance by experience in pigeons.Cah. Psychol. Cogn.18321–335.
46
WoodJ. N. (2013). Newborn chickens generate invariant object representations at the onset of visual object experience.Proc. Natl. Acad. Sci. U.S.A.11014000–14005. 10.1073/pnas.1308246110
47
WuZ.SongS.KhoslaA.YuF.ZhangL.TangX.et al (2015). “3d shapenets: A deep representation for volumetric shapes,” in Proceedings of the IEEE conference on computer vision and pattern recognition, CVPR2015, Boston, 1912–1920.
48
YaminsD. L. K.DiCarloJ. J. (2016). Using goal-driven deep learning models to understand sensory cortex.Nat. Neurosci.19356–365. 10.1038/nn.4244
49
YaminsD. L. K.HongH.CadieuC. F.SolomonE. A.SeibertD.DiCarloJ. J. (2014). Performance-optimized hierarchical models predict neural responses in higher visual cortex.Proc. Natl. Acad. Sci. U.S.A.1118619–8624. 10.1073/pnas.1403112111
50
ZhuangC.YanS.NayebiA.SchrimpfM.FrankM. C.DiCarloJ. J.et al (2021). Unsupervised neural network models of the ventral visual stream.Proc. Natl. Acad. Sci. U.S.A.118:e2014196118. 10.1073/pnas.2014196118
51
ZoccolanD.KouhM.PoggioT.DiCarloJ. J. (2007). Trade-off between object selectivity and tolerance in monkey inferotemporal cortex.J. Neurosci.2712292–12307. 10.1523/JNEUROSCI.1897-07.2007
Summary
Keywords
object detection, invariant visual perception, deep neural network, random feedforward network, learning-free model, spontaneous emergence, biologically inspired neural network, visual pathway
Citation
Cheon J, Baek S and Paik S-B (2022) Invariance of object detection in untrained deep neural networks. Front. Comput. Neurosci. 16:1030707. doi: 10.3389/fncom.2022.1030707
Received
29 August 2022
Accepted
13 October 2022
Published
03 November 2022
Volume
16 - 2022
Edited by
Chang-Eop Kim, Gachon University, South Korea
Reviewed by
Jeong-woo Sohn, Catholic Kwandong University, South Korea; Yongseok Yoo, Soongsil University, South Korea
Updates

Check for updates
Copyright
© 2022 Cheon, Baek and Paik.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Se-Bum Paik, sbpaik@kaist.ac.kr
†These authors have contributed equally to this work
This article was submitted to Frontiers in Computational Neuroscience, a section of the journal Frontiers in Computational Neuroscience
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.