Look Ma, No Labels
From Robust Perception to Self-Supervised Representation Learning for Autonomous Driving
Time: Fri 2026-10-16 10.00
Location: Kollegiesalen, Teknikringen 14
Language: English
Doctoral student: Maciej Wozniak , Robotik, perception och lärande
Opponent: Ph.D. Andrei Bursuc, Valeo.ai, Paris, France
Supervisor: Professor Patric Jensfelt, Robotik, perception och lärande; Associate professor André Tiago Abelho Pereira, Tal, musik och hörsel
QC 20260923
Abstract
The autonomous driving industry is booming, and while to a side observer all the problems may seem solved, many remain. In this thesis, I focus on the "brain" of autonomous vehicles, which is the perception models.
Autonomous driving perception can be decomposed into a backbone that maps raw sensor data into a feature representation and task heads that decode that representation into detections, segmentation, occupancy, or trajectories. Every downstream task inherits the properties of that representation: if it is brittle, overfitted to the training distribution, or sensitive to sensor degradation, the whole pipeline will fail. This thesis examines what determines the quality of that representation through the backbone architecture, the supervision used to train it, and the scale and diversity of the training data. The central argument is that representation is the highest-leverage decision when data cannot be scaled freely (under sensor and cost constraints, limited annotations, or distribution shift), and that as unlabeled data scales, leverage shifts from architecture to the choice of supervision strategy.
The thesis develops this argument in three steps. First, it shows that robustness is decided inside the backbone rather than in the sensor suite or the task head. Second, it addresses cases where deployment conditions differ from training, showing that self-supervised objectives exploiting the natural correspondence between LiDAR and camera observations transfer across sensors, domains, and downstream tasks without labels. Third, it examines how representations learned on web-scale image data can be distilled into driving-specific 3D backbones.
Taken together, the results indicate that carefully designed architectures and self-supervised objectives, rather than additional annotated driving data, are the practical route to robust perception systems.