Hypovigilance Detection and Assistance to Vehicle Drivers
Résumé
Driver drowsiness is one of the main causes of road accidents. Monitoring the behavior of the driver for the detection of drowsiness is a complex problem, which involves physiological and behavioral elements. Computer vision provides the ability to monitor the person without interfering with the driving task. An accurate estimate of the driver state, can be obtained by analyzing the facial expressions, including the eye states : eyelid closeness, blinking, or gaze fixation. A driver monitoring system by analyzing the eye conditions has three basic steps : (1) face detection ; (2) eye detection and localization ; (3) recognition of the eye states (open or closed). These steps being operational under real driving conditions, must provide a highly accurate detection response. A computer vision system, dedicated to driver monitoring, uses the driver’s face as a treatment area. Such system is governed by appropriate and robust acquisition and processing techniques that ensure stable operation. In this thesis, the proposed scheme for face detection uses Gabor’s wavelets, Principal Component Analysis (PCA), to characterize the facial region with optimal data and a Support Vector Machine (SVM) classifier for the classification phase. This first step involves a new analysis strategy using image processing methods and morphological operations, This allows to recognize the exact position of the face in real conditions. Two public databases are involved in the test phase, namely the ORL face database and the CMU-MultiPIE database. We also built our own database, representing different subjects under real and uncontrolled lighting conditions, in order to evaluate the generalization performance of our approach and its robustness to ambient changes. These three databases include the most common environmental conditions in daily driving. However, degraded detection or loss of face detection is an obstacle to the overall functioning of the system, i.e., it is impossible to analyze facial features. This occurs when the driver does not maintain a frontal position to the camera. Textures information-based method has been used in this thesis, this choice was made on the Local Binary Pattern (LBP) technique, known to be highly discriminative and robust to different environmental and textures changes. In the second step, three methods are proposed, to detect the eyes in images and video sequences, obtained from a laptop with a Web camera, under different lighting conditions. This first approach, uses a spatially enhanced LBP histogram-based feature descriptor (eLBPH), the result of which is given as input to a deep learning algorithm based on recurrent neural networks (RNN), particularly the Long Short-Term Memory (LSTM) model and the SVM classifiers. The ocular region is detected successfully, with an accuracy of 98:1% in real-time video sequences, with a computation time of 0:562 seconds. However, this method may fail to correctly detect the eyes under conditions of extreme axial (horizontal or vertical) head rotation. In addition, some image textures are not well described, because of the perspective change, and the inability of the eLBPH to discriminate certain patterns in these cases. Second approach: the problems encountered in the first approach are solved by preserving the invariance of changes in the real world. In this approach we combine the Viola-Jones method for eye detection and tracking, uniform LBPs and a chi-square statistical similarity distance. This combination enhances the performance of the classic Viola-Jones detector, providing a better estimate of eye locations. It can also overcome some of the problems encountered by the first approach. The present algorithm is validated with three public databases, namely the face database (Face GI4E), the extended Yale-B database and video sequences of (GI4E Head Pose). The algorithm works, without prior detection of the face and in real lighting conditions. This algorithm locates the eyes with an accuracy of 97:35%. In the third approach, a dictionary of invariant local features, called the spatially enhanced LBP Pyramidal histogram (ePLBPH), is proposed to represent the ocular region. The ePLBPH descriptor is the core of a new algorithm called EyeLSD, which we have proposed for ocular localization and state detection (open or closed). The EyeLSD algorithm consists of three main stages, the first stage pre-processes the image by reducing noise and improving textures. The second stage integrates two classifiers, SVM and Perceptron Multilayer (MLP), for a binary classification of eye and noneye images. A series of preprocessing and post-processing steps are implemented to improve the eye detection stage. We evaluated this algorithm on three public databases, BioID, CAS PEAL-R1 and a real world eye database ZJU Eyeblink. We also acquired and annotated our own database for different eye conditions and facial expressions. The results obtained show that the EyeLSD method is effective for locating the eyes with an accuracy of 98:12% in real scenarios. The third step, which is the final stage of the EyeLSD algorithm, focuses on recognizing the state of the eyes (open or closed), establishing an effective learning strategy for interpreting the detected eye images with the descriptor Multi-TPLBP proposed. Multi-TPLBP combines LBP’s multiple resolution capability for a rich description of eye patch information (regarding both the micro- and macro-textures of the eye model). The multi-TPLBP descriptor also aims to improve the robustness of the model with different conditions of acquisition and environment. The eye state detection step has yielded promising results with an accuracy of 95:18% and can also treat a very wide range of eye appearance than other methods compared with.
Citer ce document
Accès au document
Voir sur le dépôt sourceCe document est hébergé sur son dépôt institutionnel d'origine.
Auteur(s)
Statistiques
Consultations : 3
Téléchargements : 0