Human action recognition with ResNet, Bi-LSTM and attention
The official code for our Sensors 2025 paper. 96.6 % accuracy on UCF-101, with a webcam demo.
- When
- 2025
- Origin
- PhD research, Sensors 2025 paper
- Built with
- PyTorch
- ResNet-18
- Bi-LSTM
- Multi-head attention
- Optical flow
- UCF-101
- Paper
- Sensors, 2025
A service robot should know whether the person in front of it is walking, waving or handing something over. This framework recognises actions from short video clips and is light enough to run on a mobile robot.
How it works
- Frame selection. Optical flow picks the frames with the most motion, so the model does not waste time on near-identical frames.
- Spatial features. A ResNet-18 turns each selected frame into a 512-dimensional feature vector.
- Temporal model. A bidirectional LSTM reads the sequence in both directions.
- Attention. A multi-head attention layer weights the moments that matter for the action.
Results
On UCF-101 the model reaches 96.6 % accuracy and an F1 score of 0.97. It was also validated on a mobile robot for real-time use.
The repository contains feature extraction with and without motion-based selection, training, evaluation on the full test set or a single video, and a live webcam demo.