Large language models are trained as next-token predictors, yet this training produces internal structure far richer than the generation of plausible text. The token-level probability distributions, latent behavioral associations, and text-reasoning capabilities encoded within these models constitute formally characterizable properties that can be identified, measured, and systematically exploited to solve practical problems.
This proposal examines three such properties. First, we show that the probability distribution of an autoregressive LLM serves as a rigorous security metric: by computing the exact entropy of generated text, we produce passphrases that are simultaneously secure and memorable. Second, we demonstrate that individual tokens carry latent behavioral associations that can be discovered and leveraged for reliable tool invocation without any model fine-tuning. Third, we exploit the superior text-reasoning capabilities of LLMs relative to vision-language models to construct a massive, high-quality scientific visual question answering dataset that trains VLMs to comprehend complex figures.
Jie Li is a PhD student advised by Prof. Tom Goldstein. She studies language models and their applications.

