Tóm tắt
Skeleton-based human action recognition has attracted increasing attention because it provides a concise structural representation of human motion and is less affected by environmental factors. However, for real-deployment scenarios, the skeleton data are usually extracted automatically using human pose estimation models, resulting in joint localization errors, occlusion of missing joints, and temporal inconsistencies that affect action recognition performance. Here we systematically analyze how skeleton data quality affects the performance of a state-of-the-art Transformer-based AcTv2 model used in three input configurations: ground-truth skeletons, OpenPose-extracted skeletons, and tracking-refined skeletons. Using JHMDB datasets, we show that the performance degrades profoundly on automatically extracted skeleton data, and that our tracking-based refinement improves recognition accuracy of actions with high levels of occlusion. The results also show that an imbalance in structural distribution between action groups can also lead to instability in learning motion representations.
công trình này được cấp phép theo phép Creative Commons Ghi công 4.0 Giấy phép International . p>
Bản quyền (c) 2026 Array
