Human action recognition with ResNet, Bi-LSTM and attention

The official code for our Sensors 2025 paper. 96.6 % accuracy on UCF-101, with a webcam demo.

Architecture diagram in three stages - motion-based frame selection, ResNet-18 feature extraction, and a Bi-LSTM with an attention head feeding a classifier
When
2025
Origin
PhD research, Sensors 2025 paper
Built with
  • PyTorch
  • ResNet-18
  • Bi-LSTM
  • Multi-head attention
  • Optical flow
  • UCF-101
Paper
Sensors, 2025
View the code on GitHub

A service robot should know whether the person in front of it is walking, waving or handing something over. This framework recognises actions from short video clips and is light enough to run on a mobile robot.

How it works

  1. Frame selection. Optical flow picks the frames with the most motion, so the model does not waste time on near-identical frames.
  2. Spatial features. A ResNet-18 turns each selected frame into a 512-dimensional feature vector.
  3. Temporal model. A bidirectional LSTM reads the sequence in both directions.
  4. Attention. A multi-head attention layer weights the moments that matter for the action.

Results

On UCF-101 the model reaches 96.6 % accuracy and an F1 score of 0.97. It was also validated on a mobile robot for real-time use.

The repository contains feature extraction with and without motion-based selection, training, evaluation on the full test set or a single video, and a live webcam demo.

Browse by topic