This talk presents research updates from my group’s recent work on the interpretability of language models — specifically, on 1) interpreting latent reasoning models and 2) understanding the mechanisms behind inference-time control methods for refusal.
Sarah Wiegreffe is an assistant professor in the Department of Computer Science at the University of Maryland. She works on the explainability and interpretability of deep learning systems for language, with a focus on understanding how language models make predictions in order to make them more reliable, safe, and transparent to human users. She has been honored as a 3-time Rising Star in EECS, Machine Learning, and Generative AI. She was previously a postdoc at the Allen Institute for AI and the University of Washington and, before that, received her Ph.D. and M.S. degrees from Georgia Tech.

