Feature-optimized hybrid CNN–ViT architecture for sustainable vision-based condition assessment in agriculture
Résumé
Early detection of structural and physiological changes in plants remains a difficult challenge for computer vision because of large intra-class variation and environmental noise. This paper integrates feature enhancement using Excess Green (ExG) and Excess Red (ExR) vegetation indices with feature compression using principal component analysis (PCA) and an asymmetric convolutional neural network (CNN)--Vision Transformer (ViT) fusion architecture for multi-crop plant-disease classification. Preprocessing involves extracting ExG and ExR, performing statistical normalization, and applying PCA-based feature compression to enhance discriminative ability and reduce redundant spectral information. The CNN component generates hierarchical texture encodings, while the ViT component produces self-attention encodings suited to capturing global associations. The complementary feature spaces are combined through a cross-domain fusion layer to improve representation capability. The proposed system achieves high classification accuracy (98%) and robustness across multiple crop datasets. Although edge efficiency and explainability still need to be addressed before deployment in real-world agricultural scenarios, these aspects are outlined as future work.
Citer ce document
Accès au document
Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter
Voir l'article sur le site de la revueStatistiques
Consultations : 1
Téléchargements : 0