Accès ouvert

CV classification in a big data environment using BERT and alignment learning

Article scientifique 2023 Anglais

Résumé

Abstract Today, text classification has always been a crucial discipline in the field of text data processing. However, with the advent of Big Data, text classification has reached new heights in terms of the volume of data to be processed and the complexity of the tasks. This task is of great importance in many fields, such as spam detection, sentiment analysis, document categorization, content recommendation and many more. In this work, we focus specifically on resume classification, an area that presents unique challenges due to diverse formats, ambiguous language, and variations in applicants' work experiences. We propose an innovative approach for distributed classification of CVs using contextual alignment techniques with the job offer and the pre-trained language model BERT (Bidirectional Encoder Representations from Transformers) which constitutes a major advance in this field , and using RNN (Recurrent Neural Network) to solve text classification problems. We use a distributed processing approach to leverage the parallel computing power of Spark, enabling large volumes of data to be processed efficiently. The main objective of our study is to improve the relevance of CV classification by leveraging the pre-trained language models and distributed processing power of Spark.

Citer ce document

Chafi, S., Kabil, M., Kamouss, A. (2023). CV classification in a big data environment using BERT and alignment learning. https://doi.org/10.21203/rs.3.rs-3406344/v1

Accès au document

Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter

Voir l'article sur le site de la revue

Statistiques

Consultations : 1

Téléchargements : 0