Large language models have made remarkable progress on challenging reasoning tasks by scaling computation at both
training and inference time. However, more compute does not automatically mean better reasoning: Reinforcement learning
can over-reinforce behaviors the model already favors, extended training yields diminishing returns, and longer reasoning
traces can reflect overthinking. In this talk, I will explore how understanding the structure of LLM reasoning lets us use
computation more efficiently.
First, I will examine what reinforcement learning with verifiable rewards (RLVR) actually teaches: penalizing incorrect
responses alone is surprisingly effective, often matching or surpassing PPO and GRPO across the Pass@k spectrum while
preserving diversity. Second, I will show that RLVR weight updates are dominated by a single rank-1 direction whose
magnitude grows nearly linearly, allowing us to extrapolate checkpoints that match or approach full RLVR from only the first
15–20% of training. Finally, I will introduce an efficient inference-time scaling method by identifying deep-thinking tokens,
whose predictions keep being revised in later layers and reflect more thinking efforts. Their proportion tracks accuracy more
reliably than output length, and selecting samples by it matches self-consistency at roughly half the inference cost.
Together, these results show that efficient reasoning comes not from more computation, but from understanding what RL learns,
exploiting the structure of training dynamics, and spending inference compute where it matters.
Yu Meng is an assistant professor of the Computer Science Department at the University of Virginia. His research focuses on
developing more capable, efficient, and aligned Large Language Models (LLMs). He is a recipient of the Google PhD Fellowship
(2021), the OpenAI Superalignment Fast Grant (2024), the ACM SIGKDD Dissertation Award (2024), the Amazon Research Award
(2025), the NSF CAREER Award (2026), and was named to the Forbes 30 Under 30 Asia list (2025) and AAAI New Faculty
Highlights (2026).

