What do different models have in common?
Despite being trained independently, models can develop surprisingly similar representations of the world.
My research investigates the geometric structure behind these similarities and how they can be used to connect models across modalities, even when paired data is scarce or unavailable.
Currently, I am pursuing my Ph.D. at the Computer Vision Group at TUM supervised by Prof. Daniel Cremers under the lead of Dr. Xi Wang.
Before that, I completed my M.Sc. in Mathematics in Data Science and my B.Sc. in Mathematics with a minor in Computer Science at TUM.
Models trained independently in different modalities are often similar enough to be aligned without any pairs. We align them coarsely with a single orthogonal map and show that their shared geometry predicts how well this works, also for scientific and medical modalities.
Vision-Language models need a lot of paired training data. Can we match vision and language without any supervision? Our work shows that it could be indeed feasible.
We show that existing graph neural networks struggle with graphs at different resolutions. We propose a modification of the message passing paradigm to overcome this issue.
We use Laplace approximation to learn expressive priors for neural networks. This improves the uncertainty estimation and PAC-Bayes generalization bounds.