Performance Analysis of Deep CNN Architectures for Facial Fatigue Detection: Superiority of ResNet50 over Benchmarks

Authors

  • Timothy Christyan Department of Computer Science, Universitas Pertamina
  • Ade Irawan Department of Computer Science, Universitas Pertamina

DOI:

https://doi.org/10.25299/itjrd.2026.25422

Keywords:

Facial Fatigue , Algorithm , CNN , ResNet , Model , Accuracy potential

Abstract

Increased screen time exposure to digital screens can cause fatigue and harm health. The facial fatigue detection system is one solution for monitoring the condition of digital screen users. The latest research in this field has proposed a facial fatigue detection system using a combination of the FaceNet algorithm with k-Nearest Neighbor (kNN) or multiclass Support Vector Machine (SVM), with an accuracy of 94.68%. This paper studies the use of Convolutional Neural Network (CNN) with three architectures, namely AlexNet, Visual Geometry Group (VGG), and Residual Neural Network (ResNet), to increase the accuracy of the previous research. The UTA-RLDD facial dataset, which consists of three classes, i.e., tired, alert, and non-vigilant, is used to validate the performances. Facial features
are taken from the dataset using the Haar Cascades algorithm. After preprocessing and face extraction using the Haar Cascades algorithm, a total of 3446 facial images are used for model training and evaluation. The facial features are then resized and used to train the models; splitted into around 2400 images for training, 300 images for validation, and another 300 for testing. It is found that the ResNet architecture with 50 layers (ResNet50) delivers higher accuracy on the testing set compared to other architectures, which is 95.93%. The results also show that adding architectural parameters does not always increase the accuracy.

Downloads

Download data is not yet available.

References

[1] K. Parker, J. M. Horowitz, and R. Minkin, “Covid-19 pandemic continues to reshape work in america. pew research center,” February 2022. [Online]. Available: https://www.pewresearch.org/social-trends/2022/02/16/

[2] S. Poudel, “Research report about effect of display gadgets on eyesight quality (computer vision syndrome) of m.sc.(csit) students in tribhuvan university,” International Journal of Scientific and Engineering Research, vol. 9, no. 8, pp. 22–32, August 2018.

[3] E. J. Neophytou, L. A. Manwell, and R. Eikelboom, “Effects of excessive screen time on neurodevelopment, learning, memory, mental health, and neurodegeneration: a scoping review,” International Journal of Mental Health and Addiction, vol. 19, no. 7, December 2019.

[4] NHLBI, “Reduce screen time,” February 2013. [Online]. Available: https://www.nhlbi.nih.gov/health/educational/wecan/reduce-screen-time/index.htm

[5] Z. Yan, L. Hu, H. Chen, and F. Lu, “Computer vision syndrome: A widely spreading but largely unknown epidemic among computer users,” Computers in Human Behavior, vol. 24, no. 5, pp. 2026–2042, September 2008.

[6] F. D. Adhinata, D. P. Rakhmadani, and D. Wijayanto, “Fatigue detection on face image using facenet algorithm and k-nearest neighbor classifier,” Journal of Information Systems Engineering and Business Intelligence, vol. 7, no. 1, pp. 22–30, April 2021.

[7] K. O’Shea and R. Nash, “An introduction to convolutional neural networks,” CoRR, vol. abs/1511.08458, 2015. [Online]. Available: http://arxiv.org/abs/1511.08458

[8] I. Nagappan, J. V. T. Abraham, A. Muralidhar, and V. R. Kalluri, “Real-time driver fatigue or drowsiness detection system using face image stream,” International Journal of Civil Engineering and Technology, vol. 8, no. 11, pp. 793–802, November 2017.

[9] F. Amer and M. S. H. Al-Tamimi, “Face mask detection methods and techniques: A review,” The International Journal of Nonlinear Analysis and Applications (IJNAA), vol. 13, no. 1, pp. 3811–3823, February 2022.

[10] A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classifcation with deep convolutional neural networks,” Advances in Neural Information Processing Systems, vol. 25, no. 2, january 2012.

[11] K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” 2015.

[12] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” 2015.

[13] R. Ghoddoosian, M. Galib, and V. Athitsos, “A realistic dataset and baseline temporal model for early drowsiness detection,” 2019.

[14] B. Pan, R. Panda, C. Fosco, C.-C. Lin, A. Andonian, Y. Meng, K. Saenko, A. Oliva, and R. Feris, “Va-red2: Video adaptive redundancy reduction,” 2021. [Online]. Available: https://arxiv.org/abs/2102.07887

[15] P. Viola and M. Jones, “Rapid object detection using a boosted cascade of simple features,” vol. 1, 02 2001, pp. I–511.

[16] G. E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R. R. Salakhutdinov, “Improving neural networks by preventing co-adaptation of feature detectors,” 2012. [Online]. Available: https://arxiv.org/abs/1207.0580

[17] D. C. Cires¸an, U. Meier, J. Masci, L. M. Gambardella, and J. Schmidhuber, “High-performance neural networks for visual object classification,” 2011. [Online]. Available: https://arxiv.org/abs/1102.0183

[18] G. E. Dahl, T. N. Sainath, and G. E. Hinton, “Improving deep neural networks for lvcsr using rectifed linear units and drop-out,” 2013 IEEE international conference on acoustics, speech and signal processing, May 2013.

[19] R. Pascanu, T. Mikolov, and Y. Bengio, “On the difficulty of training recurrent neural networks,” 2013. [Online]. Available: https://arxiv.org/abs/1211.5063

[20] D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” 2017. [Online]. Available: https://arxiv.org/abs/1412.6980

Downloads

Published

2026-09-30

How to Cite

Christyan, T., & Irawan, A. (2026). Performance Analysis of Deep CNN Architectures for Facial Fatigue Detection: Superiority of ResNet50 over Benchmarks. IT Journal Research and Development, 11(1), 37–54. https://doi.org/10.25299/itjrd.2026.25422

Issue

Section

Articles