TrajMamba: predicting where people will walk
A 262 K-parameter Mamba model that predicts the next three seconds of a person’s path from an RGB-D camera, in real time on a Jetson Orin.
- When
- 2026
- Origin
- PhD research, paper accepted at ICCMA 2026
- Built with
- ROS 2 Humble and Jazzy
- PyTorch
- Mamba
- MediaPipe Pose
- Jetson Orin
- TIAGo
A robot that plans around people needs to know where they will be, not only where they are. This ROS 2 package tracks up to four people and predicts each one’s path for the next three seconds.
Pipeline
- MediaPipe Pose finds the hip centre of each person in the RGB image.
- The depth image and camera intrinsics turn that pixel into a 3D point.
- A nearest-neighbour tracker links detections over time.
- After about two seconds of observation, TrajMamba predicts the next three.
TrajMamba is a bidirectional Mamba (selective state space) encoder with a GRU decoder. It has about 262 K parameters, small enough to run on board an NVIDIA Jetson Orin beside the rest of the navigation stack.
Output
The node publishes an annotated image, RViz markers and a PoseArray of predicted paths, so a costmap plugin or planner can consume them directly. A live window shows skeletons, the observed trail, the prediction with its uncertainty cone and a depth gauge.
Known limit
Each person is predicted independently. Interaction between people is not modelled yet, and that is what I am working on next.
The work was supported by BMFTR (Edison) and the European Regional Development Fund (ENABLING, ORAKEL).