log in  |  register  |  feedback?  |  help  |  web accessibility
PhD Proposal: Exploiting Formally Characterizable Properties of Large Language Models
Jie Li
IRB-3137 https://umd.zoom.us/j/96024552819?pwd=Z77D0rvy5OmBfUqnfK52pFvVaOi6r4.1
Monday, July 27, 2026, 12:30-2:00 pm
  • You are subscribed to this talk through .
  • You are watching this talk through .
  • You are subscribed to this talk. (unsubscribe, watch)
  • You are watching this talk. (unwatch, subscribe)
  • You are not subscribed to this talk. (watch, subscribe)
Abstract

Large language models are trained as next-token predictors, yet this training produces internal structure far richer than the generation of plausible text. The token-level probability distributions, latent behavioral associations, and text-reasoning capabilities encoded within these models constitute formally characterizable properties that can be identified, measured, and systematically exploited to solve practical problems.

This proposal examines three such properties. First, we show that the probability distribution of an autoregressive LLM serves as a rigorous security metric: by computing the exact entropy of generated text, we produce passphrases that are simultaneously secure and memorable. Second, we demonstrate that individual tokens carry latent behavioral associations that can be discovered and leveraged for reliable tool invocation without any model fine-tuning. Third, we exploit the superior text-reasoning capabilities of LLMs relative to vision-language models to construct a massive, high-quality scientific visual question answering dataset that trains VLMs to comprehend complex figures.

Bio

Jie Li is a PhD student advised by Prof. Tom Goldstein. She studies language models and their applications.

This talk is organized by Migo Gui