Aleksandra Bakalova
Saarbrücken, Germany
I am Sasha (short for Aleksandra) Bakalova, PhD student in LaCoCo Lab at Saarland University supervised by Prof. Michael Hahn. I am also a member of RTG ‘‘Neuroexplicit Models of Language, Vision, and Action’’.
I did my Masters in Language Science and Technology at Saarland University, and my undergraduate in Applied Mathematics and Computer Science at HSE University. I learned a lot about deep learning at Yandex School of Data Analysis.
I want to develop interpretability methods that are faithful to models’ internal computations, and use these methods to understand what solutions models learn, and what drives them toward these specific solutions. I say that I work at the intersection of mechanistic interpretability and deep learning theory, and I think of AI safety and further development of theory as the closest applications of my research. Currently, I am interested in finding and interpreting generalizable algorithms learned by models as correctly as possible. For more details, take a look at the blog post about our recent paper, where we extract programs from length-generalizable transformers trained on simple tasks.
If you wish to chat, feel free to drop me an email!
News
| May 15, 2026 | Our paper on mechanistic understanding of how LLMs learn in-context when some of the demostrations are ambiguious was accepted to ICML 2026! Read the paper on arxiv. |
|---|---|
| Mar 04, 2026 | We propose a way to rewrite trained models into D-RASP – a programming language that mimics transformers, and show that it works, at least in the controlled setting of small models and algorithmic and formal languages tasks. See the paper on arxiv and this blog post for more visualizations and examples of decompiled programs. UPD: accepted to ICML 2026! |
| Nov 01, 2025 | Excited to start a new chapter as a PhD student at the RTG ‘‘Neuroexplicit Models of Language, Vision, and Action’’ under the supervision of Prof. Michael Hahn! |
| Oct 15, 2025 | Born a Transformer – Always a Transformer? is accepted to Neurips 2025! |
| Jul 31, 2025 | Attended EEML 2025, got a best poster award! Walking around Sarajevo, getting inspired by the amazing people there and their work: it was a lot of fun. |
| Jul 07, 2025 | Our work on understanding in-context learning is accepted to COLM 2025! |
| May 30, 2025 | Born a Transformer – Always a Transformer?: new preprint is on arxiv! We try to bring theory closer to practice and answer a question: how do theoretical limitations of transformers manifest in pre-trained models? See Yana’s X post for a short explanation. |
| Mar 31, 2025 | Our preprint on understanding in-context learning in Gemma2-2b is on arxiv! We found a circuit in the model that performs in-context learning and interpreted the information flow in it. See this X post for a short explanation. |
| Aug 01, 2024 | Started my Erasmus semester at Charles University in Prague. In love with the city! |
| Apr 30, 2024 | Attended ALPS 2024 for some inspiration and breathtaking views :) |
| Oct 30, 2023 | Started a job as a Research Assistant at LaCoCo Lab with Dr. Michael Hahn. Very excited to work on LLM interpretability with such amazing colleagues! |
| Oct 01, 2023 | Started my Masters in Language Science and Technology at Saarland University. |
Publications
2026
- How Few-Shot Examples Add Up: A Causal Decomposition of Function Vectors in In-Context Learning2026
-
2025
- Born a Transformer–Always a Transformer? On the Effect of Pretraining on Architectural AbilitiesThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025
- Contextualize-then-Aggregate: Circuits for In-Context Learning in Gemma-2 2BSecond Conference on Language Modeling, 2025