Skip to main content
To KTH's start page

Temporal Modeling of Dynamic 3D Scenes for Self-supervised Scene Flow

Time: Fri 2026-10-23 09.00

Location: Kollegiesalen, Brinellvägen 8, Stockholm

Video link: https://kth-se.zoom.us/j/68797563146

Language: English

Subject area: Computer Science

Doctoral student: Qingwen Zhang , Robotik, perception och lärande

Opponent: Professor Teresa Vidal Calleja, University of Technology Sydney, Sydney, Australia

Supervisor: Professor Patric Jensfelt, Robotik, perception och lärande; Assistant Professor Olov Andersson, Robotik, perception och lärande

Export to calendar

QC 20260928

Abstract

The physical world is never still. This thesis studies how the 3D motion of a dynamic scene can be estimated from point clouds without human annotation. Scene flow assigns a displacement vector to every observed point, but dense motion labels are prohibitively expensive, which has kept the representation from practical use.

We approach the problem through the physical nature of motion. Unlike the categories behind detection or segmentation, motion is not defined by people; it is a physical quantity of the scene, and three of its properties are exploited here. It is measurable from geometry, so supervision can be derived from observations rather than annotators. It is continuous in time, so past frames constrain the present one. And it depends on kinematics rather than appearance, so it can be generated at scale in simulation and transferred to real sensors.

We first show that feed-forward models trained with point-wise matching systematically underestimate dynamic motion, because nearest-neighbor correspondences on large moving objects point to the wrong surface. Separating static structure from dynamic objects and enforcing object-level kinematic constraints yields reliable label-free supervision from a pair of frames. To exploit the temporal continuity of motion, we then extend supervision to multiple frames, which raises two problems: computation that grows with the temporal window, which we keep constant through a compact residual representation, and correspondences that shift abruptly across frames, which we resolve by aggregating the most temporally consistent cues from a candidate pool, keeping multi-frame self-supervision stable. Even with such carefully designed objectives, self-supervision still relies on proxy signals derived from geometric assumptions. These signals remain noisy under sensor sparsity and occlusion, so scaling unlabeled real data yields diminishing returns. Because motion depends on kinematics rather than appearance, we instead turn to simulation, prioritizing motion complexity over sensor realism to generate large-scale synthetic data with noise-free labels derived directly from simulator states. The resulting motion prior generalizes zero-shot across sensors and sharply reduces the need for real annotations. Finally, we show that reliable flow estimates act as point-level velocity fields that remove motion distortion from raw LiDAR sweeps.

Together, these results show that the physical objectivity of motion is more than a conceptual idea: it is a property that can be exploited to make dynamic scene understanding scalable and generalizable without human supervision.

Link to DiVA