TrajMamba: predicting where people will walk

A 262 K-parameter Mamba model that predicts the next three seconds of a person’s path from an RGB-D camera, in real time on a Jetson Orin.

Camera view from a mobile robot showing a tracked person with a skeleton overlay, the observed walking path in blue and the predicted path in green
When
2026
Origin
PhD research, paper accepted at ICCMA 2026
Built with
  • ROS 2 Humble and Jazzy
  • PyTorch
  • Mamba
  • MediaPipe Pose
  • Jetson Orin
  • TIAGo
View the code on GitHub

A robot that plans around people needs to know where they will be, not only where they are. This ROS 2 package tracks up to four people and predicts each one’s path for the next three seconds.

Pipeline

  1. MediaPipe Pose finds the hip centre of each person in the RGB image.
  2. The depth image and camera intrinsics turn that pixel into a 3D point.
  3. A nearest-neighbour tracker links detections over time.
  4. After about two seconds of observation, TrajMamba predicts the next three.

TrajMamba is a bidirectional Mamba (selective state space) encoder with a GRU decoder. It has about 262 K parameters, small enough to run on board an NVIDIA Jetson Orin beside the rest of the navigation stack.

Output

The node publishes an annotated image, RViz markers and a PoseArray of predicted paths, so a costmap plugin or planner can consume them directly. A live window shows skeletons, the observed trail, the prediction with its uncertainty cone and a depth gauge.

Known limit

Each person is predicted independently. Interaction between people is not modelled yet, and that is what I am working on next.

The work was supported by BMFTR (Edison) and the European Regional Development Fund (ENABLING, ORAKEL).

Browse by topic