COATRES-NET: CNN COM ATENÇÃO PARA RECONHECIMENTO DE GESTOS

Authors

DOI:

https://doi.org/10.18623/rvd.v23.7878

Keywords:

Aprendizado Profundo, Mecanismos de Atenção, Reabilitação Neuromotora Domiciliar, Reconhecimento de Gestos de Mão, Redes Neurais Convolucionais

Abstract

Sistemas de reabilitação neuromotora domiciliar baseados em visão computacional dependem de reconhecimento gestual preciso e computacionalmente leve, mas a literatura carece de comparações sistemáticas entre arquiteturas candidatas sob protocolo estatístico rigoroso. Este trabalho apresenta a CoAtRes-Net, rede convolucional autoral que combina atenção de canal e de coordenada com conexões residuais, aplicada ao reconhecimento de nove gestos de mão a partir de landmarks esparsos do MediaPipe Hands. A arquitetura foi comparada estatisticamente a seis arquiteturas representativas, incluindo três backbones pré-treinados, sob validação cruzada estratificada de cinco dobras, conjunto de teste independente e testes pareados com correção para comparações múltiplas. A CoAtRes-Net obteve o melhor desempenho entre as sete arquiteturas avaliadas, com F1-macro de 84,9% na validação cruzada e 84,3% no teste independente, além de latência de 5,56 ms em CPU convencional, sem GPU dedicada. Uma ablação fatorial indica que a conexão residual, não a atenção de coordenada, explica a maior parte desse ganho. Quatro tentativas de refinamento arquitetural foram testadas e refutadas com o mesmo rigor estatístico. Os resultados indicam que a arquitetura é um componente promissor para reconhecimento gestual de baixo custo computacional, embora sua aplicação clínica dependa de validações futuras não conduzidas aqui.

References

ALONAZI, M. et al. Smart Healthcare Hand Gesture Recognition Using CNN-Based Detector and Deep Belief Network. IEEE Access, v. 11, p. 84922–84933, 2023. DOI: 10.1109/ACCESS.2023.3289389.

AMPRIMO, G. et al. Hand tracking for clinical applications: validation of the Google MediaPipe Hand (GMH) and the depth-enhanced GMH-D frameworks. Biomedical Signal Processing and Control, v. 96, p. 106508, 2024. DOI: 10.1016/j.bspc.2024.106508.

BOUDREAULT-MORALES, G. E. et al. The effect of depth data and upper limb impairment on lightweight monocular RGB human pose estimation models. BioMedical Engineering OnLine, v. 24, p. 12, 2025. DOI: 10.1186/s12938-025-01347-y.

BRASIL. Lei nº 13.709, de 14 de agosto de 2018. Lei Geral de Proteção de Dados Pessoais (LGPD). Diário Oficial da União, Brasília, DF, 15 ago. 2018. Disponível em: https://www.planalto.gov.br/ccivil_03/_ato2015-2018/2018/lei/l13709.htm. Acesso em: 05 ago. 2026.

CHENG, Q. et al. A Dense-Sparse Complementary Network for Human Action Recognition based on RGB and Skeleton Modalities. Expert Systems with Applications, v. 244, p. 123061, 2024.

CUI, H. et al. Dstsa-gcn: Advancing skeleton-based gesture recognition with semantic-aware spatio-temporal topology modeling. Neurocomputing, 2025.

CUNHA, B.; MAÇÃES, J.; AMORIM, I. Smartphone-Based Markerless Motion Capture for Accessible Rehabilitation: A Computer Vision Study. Sensors, v. 25, n. 17, p. 5428, 2025. DOI: 10.3390/s25175428.

DEMŠAR, J. Statistical Comparisons of Classifiers over Multiple Data Sets. Journal of Machine Learning Research, v. 7, p. 1–30, 2006.

DOSOVITSKIY, A. et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. International Conference on Learning Representations (ICLR), 2021. arXiv:2010.11929.

DUAN, H. et al. Revisiting Skeleton-Based Action Recognition. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 2969–2978, 2022.

FERNANDES, H. S.; LIBERATO, I. B. Avaliação da performance de alunos com limitações sem atividades físico-motoras através da inteligência artificial em conjunto com o Kinect [Internet]. Repositório Institucional IFES, 2021. Disponível em: https://repositorio.ifes.edu.br/handle/123456789/1605. Acesso em: 20 mar. 2025.

HADJIOSIF, A. M. et al. Tiny visual latencies can profoundly impair implicit sensorimotor learning. Scientific Reports, 2025.

HE, K. et al. Deep Residual Learning for Image Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), p. 770–778, 2016.

HOLM, S. A Simple Sequentially Rejective Multiple Test Procedure. Scandinavian Journal of Statistics, v. 6, n. 2, p. 65–70, 1979.

HOU, Q.; ZHOU, D.; FENG, J. Coordinate Attention for Efficient Mobile Network Design. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021.

HU, R. et al. Effective evaluation of HGcnMLP method for markerless 3D pose estimation of musculoskeletal diseases patients based on smartphone monocular video. Frontiers in Bioengineering and Biotechnology, v. 12, p. 1335251, 2024. DOI: 10.3389/fbioe.2024.1335251.

HU, J.; SHEN, L.; SUN, G. Squeeze-and-Excitation Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), p. 7132–7141, 2018.

IOFFE, S.; SZEGEDY, C. Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. Proceedings of the 32nd International Conference on Machine Learning (ICML), p. 448–456, 2015.

KE, Q. et al. A New Representation of Skeleton Sequences for 3D Action Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), p. 4570–4579, 2017. DOI: 10.1109/CVPR.2017.486.

LATRECHE, A. et al. Reliability and validity analysis of MediaPipe-based measurement system for some human rehabilitation motions. Measurement, v. 214, p. 112826, 2023.

LI, J.; AL-QANESS, M. A. A. SA-STGCN: Structural-Adaptive Spatio-Temporal Graph Convolution with Spatio-Temporal Attunement for skeleton-based gesture recognition. Robotics and Autonomous Systems, 2026.

LIN, J.; GAN, C.; HAN, S. TSM: Temporal Shift Module for Efficient Video Understanding. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), p. 7083–7093, 2019.

LIU, J. et al. Temporal decoupling graph convolutional network for skeleton-based gesture recognition. IEEE Transactions on Multimedia, v. 26, p. 811–823, 2024.

LIU, Z. et al. A ConvNet for the 2020s. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 11976–11986, 2022.

LIU, Z. et al. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), p. 9992–10002, 2021.

LOSHCHILOV, I.; HUTTER, F. Decoupled Weight Decay Regularization. International Conference on Learning Representations (ICLR), 2019. arXiv:1711.05101.

LOSHCHILOV, I.; HUTTER, F. SGDR: Stochastic Gradient Descent with Warm Restarts. International Conference on Learning Representations (ICLR), 2017.

LUGARESI, C. et al. MediaPipe: A Framework for Building Perception Pipelines. arXiv preprint, 2019. arXiv:1906.08172.

MATERZYNSKA, J. et al. The Jester Dataset: A Large-Scale Video Dataset of Human Gestures. Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2019.

OGG, A. D. et al. Integração de Visão Computacional e Inteligência Artificial na Fisioterapia Domiciliar. ISLA 2024 Proceedings, 2024. Disponível em: https://aisel.aisnet.org/isla2024/12. Acesso em: 14 jan. 2025.

SRIVASTAVA, N. et al. Dropout: A Simple Way to Prevent Neural Networks from Overfitting. Journal of Machine Learning Research, v. 15, n. 1, p. 1929–1958, 2014.

SZEGEDY, C. et al. Rethinking the Inception Architecture for Computer Vision. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), p. 2818–2826, 2016.

TAN, M.; LE, Q. V. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. Proceedings of the 36th International Conference on Machine Learning (ICML), p. 6105–6114, 2019.

TSINGANOS, P. et al. Real-Time Analysis of Hand Gesture Recognition with Temporal Convolutional Networks. Sensors, v. 22, n. 5, p. 1694, 2022. DOI: 10.3390/s22051694.

VASWANI, A. et al. Attention Is All You Need. Advances in Neural Information Processing Systems (NeurIPS), p. 6000–6010, 2017.

WAGH, V.; SCOTT, M. W.; KRAEUTNER, S. N. Quantifying Similarities Between MediaPipe and a Known Standard to Address Issues in Tracking 2D Upper Limb Trajectories: Proof of Concept Study. JMIR Form Res, v. 8, 2024. DOI: 10.2196/56682.

WANG, Q. et al. ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 11531–11539, 2020.

WANG, X. et al. Non-Local Neural Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), p. 7794–7803, 2018.

WILCOXON, F. Individual Comparisons by Ranking Methods. Biometrics Bulletin, v. 1, n. 6, p. 80–83, 1945. DOI: 10.2307/3001968.

WOO, S. et al. CBAM: Convolutional Block Attention Module. Proceedings of the European Conference on Computer Vision (ECCV), p. 3–19, 2018.

YAN, S.; XIONG, Y.; LIN, D. Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition. Proceedings of the AAAI Conference on Artificial Intelligence, v. 32, n. 1, 2018. DOI: 10.1609/aaai.v32i1.12328.

ZAHER, M. et al. Fusing CNNs and attention-mechanisms to improve real-time indoor Human Activity Recognition for classifying home-based physical rehabilitation exercises. Computers in Biology and Medicine, v. 184, p. 109399, 2025. DOI: 10.1016/j.compbiomed.2024.109399.

ZAINUDDIN, A. A. et al. Calibrating Hand Gesture Recognition for Stroke Rehabilitation Internet-of-Things (RIOT) Using MediaPipe in Smart Healthcare Systems. International Journal of Advanced Computer Science and Applications (IJACSA), v. 15, n. 7, 2024.

ZHANG, F. et al. MediaPipe Hands: On-device Real-time Hand Tracking. arXiv preprint, 2020. arXiv:2006.10214.

ZHONG, E. et al. Real-Time Monocular Skeleton-Based Hand Gesture Recognition Using 3D-Jointsformer. Sensors, v. 23, n. 16, 2023. DOI: 10.3390/s23167066.

Published

2026-08-18

How to Cite

Amorim, D. G. P., Souza, V. P. M. de, Santos, W. J. C. dos, Coutinho, R. M. P., Seixas, V. N. C., Silva, E. M. da, … Conte, T. N. M. de S. (2026). COATRES-NET: CNN COM ATENÇÃO PARA RECONHECIMENTO DE GESTOS. Veredas Do Direito, 23(14), e237878. https://doi.org/10.18623/rvd.v23.7878