Promoting the Shift From Pixel-Level Correlations to Object Semantics Learning by Rethinking Computer Vision Benchmark Data Sets-Reference-Cited by-同舟云学术

Promoting the Shift From Pixel-Level Correlations to Object Semantics Learning by Rethinking Computer Vision Benchmark Data Sets

Published:2024-07-19 Issue:8 Volume:36 Page:1626-1642
ISSN:0899-7667
Container-title:Neural Computation
language:en
Short-container-title:

Author:

Osório Maria¹,Wichert Andreas²

Affiliation:

1. Department of Computer Science and Engineering, INESC-ID and Instituto Superior Técnico, University of Lisbon, 2744-016 Porto Salvo, Portugal maria.osorio@tecnico.ulisboa.pt

2. Department of Computer Science and Engineering, INESC-ID and Instituto Superior Técnico, University of Lisbon, 2744-016 Porto Salvo, Portugal andreas.wichert@tecnico.ulisboa.pt

Abstract

Abstract In computer vision research, convolutional neural networks (CNNs) have demonstrated remarkable capabilities at extracting patterns from raw pixel data, achieving state-of-the-art recognition accuracy. However, they significantly differ from human visual perception, prioritizing pixel-level correlations and statistical patterns, often overlooking object semantics. To explore this difference, we propose an approach that isolates core visual features crucial for human perception and object recognition: color, texture, and shape. In experiments on three benchmarks—Fruits 360, CIFAR-10, and Fashion MNIST—each visual feature is individually input into a neural network. Results reveal data set–dependent variations in classification accuracy, highlighting that deep learning models tend to learn pixel-level correlations instead of fundamental visual features. To validate this observation, we used various combinations of concatenated visual features as input for a neural network on the CIFAR-10 data set. CNNs excel at learning statistical patterns in images, achieving exceptional performance when training and test data share similar distributions. To substantiate this point, we trained a CNN on CIFAR-10 data set and evaluated its performance on the “dog” class from CIFAR-10 and on an equivalent number of examples from the Stanford Dogs data set. The CNN poor performance on Stanford Dogs images underlines the disparity between deep learning and human visual perception, highlighting the need for models that learn object semantics. Specialized benchmark data sets with controlled variations hold promise for aligning learned representations with human cognition in computer vision research.

Publisher

MIT Press

Link

https://direct.mit.edu/neco/article-pdf/36/8/1626/2462307/neco_a_01677.pdf

Reference30 articles.

1. Multi-resolution intrinsic texture geometry-based local binary pattern for texture classification;Alpaslan;IEEE Access,2020

2. Neuroscience needs network science;Barabási;Journal of Neuroscience,2023

3. A computational approach to edge detection;Canny;IEEE Transactions on Pattern Analysis and Machine Intelligence,1986

4. Separate processing of texture and form in the ventral stream: Evidence from FMRI and visual agnosia;Cavina-Pratesi;Cerebral Cortex,2010