Foundation models are becoming runtime decision-makers. At inference time, they decode under constraints, compare alternatives, search over hypotheses, call tools, coordinate with other agents, verify intermediate states, and defend against adversarial inputs. This creates a central control problem: how should an agent choose its next computation, action, or workflow under uncertainty, risk, latency, and safety constraints? In this talk, I will present adaptive test-time control as a unifying framework for reasoning and planning agents. Rather than treating decoding, search, reflection, collaboration, verification, and escalation as fixed inference recipes, we view them as controllable actions selected by a runtime policy. I will ground this framework in recent work from our group across a hierarchy of test-time control. At the generation level, GenARM, Transfer Q-Star, and Collab show how reward models, value estimates, and mixtures of agents can steer decoding and alignment at inference time. But control is only as reliable as its critic; ReForm closes this loop by using reward-guided failure discovery to expose and patch reward-model errors. At the action level, Agentic Critical Training trains agents to judge better actions among alternatives, turning reflection into action-quality control rather than imitation. At the workflow level, FlowBank adaptively selects among complementary multi-agent workflows under performance–cost tradeoffs. Together, these works suggest a path from controlled decoding to controlled agency: agents that allocate computation, evidence, collaboration, and safety intervention where they matter most. The broader message is that reliable planning agents require runtime policies for allocating computation, evidence, collaboration, verification, workflow structure, and safety intervention. The goal is to build agents that are selective, robust, and accountable — systems that deliver reliable behavior per unit compute in interactive, multimodal, and tool-rich environments.
[2] FlowBank: Query-Adaptive Agentic Workflows Optimization through Precompute-and-Reuse
Furong Huang is an Associate Professor in the Department of Computer Science at the University of Maryland. She is also affiliated with the Center for Machine Learning at Institute for Advanced Computer Studies, the Maryland Robotics Center, the Applied Mathematics, Statistics, and Scientific Computation Program, and the Department of Electrical and Computer Engineering.
Her research bridges trustworthy machine learning, sequential decision-making, and generative AI, with a strong emphasis on developing foundation models for robotics. These models aim to unify perception, planning, and control across diverse robotic platforms. Dr. Huang’s broader vision is to build intelligent systems that are not only high-performing, but also reliable, interpretable, and aligned with human values. Her work combines theoretical rigor with real-world impact, enabling systems that are robust, adaptable, and safe.

