Robots that move among people

My PhD research follows the data through a mobile robot: from raw sensor readings to a map, from the map to an understanding of the people in it, and from there to motion that people find acceptable. Everything is tested on a PAL Robotics TIAGo with ROS 2 and Nav2.

  1. 1SenseMulti-sensor fusion for 2D mapping
  2. 2MapVisual SLAM in dynamic scenes
  3. 3UnderstandHuman action, trajectory and engagement
  4. 4NavigateSocially-aware navigation
  5. 5ReasonLLMs for robot task control

1Sense

Fusing cameras and LiDAR into one reliable scan

A 2D LiDAR sees a single slice of the room. It misses table tops, chair legs and anything above or below its plane. I project the point clouds of four RGB-D cameras into that plane, filter the noise and merge them with two LiDAR scans. The mapping algorithm, Enhanced Gmapping (EGM), adds adaptive particle resampling and handles the cases where a particle filter normally degenerates.

Localisation drift fell from 1.56 m to 0.85 m and the robot reached 95 % of its goals indoors.

Particle-filter localisation, the idea behind Gmapping. Each dot is a guess of where the robot is. Guesses that explain the laser scan survive resampling. Drag the robot to kidnap it and watch the filter recover.

2Map

Keeping the map correct while the scene moves

Classic visual SLAM assumes a static world. When a person walks through the image, features on that person drag the estimated camera pose with them. SAR-SLAM segments likely movers with YOLOv8, checks with a RANSAC homography whether they are really moving, and fuses both answers before features reach ORB-SLAM3. A person standing still keeps their features; a walking person loses them. It plugs in without changing the ORB-SLAM3 code.

Up to 96 % lower absolute trajectory error than ORB-SLAM3, and best or joint-best on all eight TUM RGB-D sequences tested.

A camera view with tracked features. With the filter on, features on the walking person are rejected while the person standing still keeps theirs, and the pose error stays flat. The demo switches the filter with each pass until you take over.

3Understand

Recognising what people do and where they will go

A robot that shares space with people needs to know more than where they are. I work on three signals. Action recognition tells the robot what a person is doing, with models small enough for an embedded GPU. Trajectory prediction with TrajMamba, a 262 K-parameter bidirectional Mamba model, estimates the next three seconds of a person’s path from RGB-D. Engagement and gaze estimation, with colleagues, tell the robot whether someone wants to interact at all.

96.6 % accuracy on UCF-101, and REMA-Net reaches 96.8 % at 11.82 ms per inference on a Jetson Nano.

Camera view from a TIAGo robot with a tracked skeleton, the observed walking path and the predicted path of a person drawn on the floor
TrajMamba running on the robot. Blue is the observed path, green the predicted next three seconds.

5Reason

Letting language models decide what to do next

With colleagues I studied how large language models can be wired into a robot: open-loop, where the model writes a plan once; closed-loop, where it sees the result of each step; and fully LLM-driven control. We tested the taxonomy on TIAGo and in simulation with GPT-4 and Vicuna.

Closed-loop systems handled unexpected situations effectively, with human-like problem solving.

Papers

Platform and funding

All methods are validated in simulation and on real hardware: a PAL Robotics TIAGo with ROS 2 Humble and Nav2, and Jetson Nano and Orin boards. The work is supervised by Prof. Dr. Ayoub Al-Hamadi in the Neuro-Information Technology (NIT) group and has been funded through these projects.

  • BMBF AutoKoWAT-3DMAt
  • DFG SEMIAC
  • EFRE ENABLING
  • EFRE RoboLab
  • BMFTR Edison
  • ERDF ORAKEL

Browse by topic