Data, Geometry and Homology
Time: Wed 2026-09-23 10.00
Location: F3 (Flodis), Lindstedtsvägen 26 & 28, Stockholm
Language: English
Subject area: Applied and Computational Mathematics
Doctoral student: Jens Agerberg , Algebra, kombinatorik och topologi
Opponent: Professor Bastian Grossenbacher-Rieck, University of Fribourg, Av. de l'Europe 20, CH-1700 Fribourg, Switzerland
Supervisor: Associate Professor Martina Scolamiero, Algebra, kombinatorik och topologi; Professor Wojciech Chachólski, Algebra, kombinatorik och topologi
QC 2026-09-01
Abstract
Modern datasets are increasingly complex and heterogeneous. Analyzing such data raises fundamental questions about the mathematical spaces in which data should be represented, how objects in these spaces should be compared, and which computable invariants preserve information relevant to a given task. This thesis studies these questions with topological data analysis as a central framework, using homology as a language for describing the geometry of data.
The first part of the thesis develops distances and invariants for spaces of persistence modules. The starting point is categorical: we regard data as objects living in categories equipped with enough algebraic structure to define a notion of size. In particular, abelian categories provide kernels, cokernels, and exact sequences, allowing distances between objects to be constructed from the failure of morphisms to be isomorphisms. Persistence modules form a central example: they encode how homological features appear and disappear along a filtration, and may be viewed as functors from a partially ordered parameter space to vector spaces. Within this framework, we study distances induced by contours, which provide a flexible way of specifying the geometry of the parameter space. This setting leads to compactness results for families of multidimensional persistence modules. In the one-dimensional setting, where a barcode decomposition is available, we develop algebraic Wasserstein distances based on ℓp norms of contour-dependent bar lifetimes. Using these distances, we define Wasserstein stable ranks, stable and computable invariants whose interpretable parameters can be learned for a given task.
The second part of the thesis moves from the mathematical framework to applications in neuroscience, where cellular morphologies provide natural examples of structured geometric data. Microglia and other branched cells can be represented as rooted trees embedded in three-dimensional space, and their morphology can be characterized using topological morphology descriptors. In the morphOMICs pipeline, such descriptors are combined with vectorizations, bootstrapping, dimensionality reduction, and classification in order to map microglial morphology across brain regions and sexes, and through development, disease progression, and experimental perturbations. This gives a data-driven atlas of microglial morphology that avoids relying on preselected scalar morphometric features.
We further introduce the chromatic topological morphology descriptor (chromatic TMD) to study intracellular organization in branched cells. Here a microglial cell is represented by a rooted tree, while CD68-positive and mitochondria organelles are represented by subgraphs of that tree. The inclusion of the organelle subgraph into the cell tree induces a morphism of persistence modules, and the image, kernel, and cokernel of this morphism describe complementary aspects of organelle organization: where organelles occupy branches, where they co-localize within branch structures, and where they are absent. An efficient tree-based algorithm is developed for computing these descriptors. Applied to retinal microglia, the method reveals organelle-specific spatial programs: CD68-positive organelles reorganize in a layer- and injury-dependent manner, while mitochondrial organization remains more closely coupled to the underlying branching morphology.
The third part of the thesis studies how stable homological invariants can be used in machine learning. Stable ranks provide a bridge from persistence modules to function spaces or finite-dimensional vector spaces, making persistence-based information accessible to kernel methods and neural networks. We introduce stable rank kernels, in which the choice of distance on persistence modules determines the stable rank and, consequently, the similarities encoded by the kernel. Varying this distance through contours can improve supervised learning performance. We also study subsampling-based stable ranks, in which probability distributions on a reference dataset are used to draw many subsamples, compute persistent homology and the corresponding stable ranks, and average the resulting functions. Different choices of distribution yield global descriptors of datasets or relative descriptors of points in the ambient space with respect to a reference object.
Finally, we investigate robustness in persistence-based learning. Persistent homology is stable with respect to suitable metrics, but these guarantees need not be preserved when persistence modules are processed by neural networks. We therefore introduce a stable rank network, combining stable rank vectorizations with Lipschitz neural network layers. This architecture has a controlled Lipschitz constant and yields sample-wise certificates of robustness in Wasserstein or bottleneck distance. This shows that topological stability can be preserved through a learning pipeline and used to certify robustness against adversarial perturbations.