log in  |  register  |  feedback?  |  help  |  web accessibility
PhD Proposal: Music as Structured Prior over Human Motion: Inference, Generation, and Control
Seong Yoo
IRB 4107 https://umd.zoom.us/j/4523376959?pwd=dk42ZFpzY0psWGRRK0hCclQ0NjVOUT09&omn=98400435203
Thursday, August 27, 2026, 10:00-11:30 am
  • You are subscribed to this talk through .
  • You are watching this talk through .
  • You are subscribed to this talk. (unsubscribe, watch)
  • You are watching this talk. (unwatch, subscribe)
  • You are not subscribed to this talk. (watch, subscribe)
Abstract

One of the unique properties of humans is our musical behavior. Group dancing is a universal phenomenon across different cultures, ethnicities, and eras. We have a strong innate desire to move our bodies and synchronize with others through musical activity. Simple beats make people dance, and songs let us share emotion and transmit ancestral knowledge across generations. Understanding human behavior in musical contexts bridges the gap between our instinct and intelligence. This thesis argues that music is not merely a conditioning signal for models of human motion, but a structured prior over body kinematics, a temporally organized constraint on which movements are plausible at each instant. I make this claim at three levels: Inference, Generation, and Control.

Inference. Visual data captures low-frequency motion well but often fails to capture high-frequency detail due to occlusion, pixel saturation, and low sampling rates. Audio signals carry exactly that missing detail, yet they are inefficient for estimating low-frequency information. VioPose exploits this complementarity through hierarchical audiovisual inference, using audio to compensate for the limitations of visual data and recover 4D pose in violin playing scenarios.

Generation. Existing literature on dance generation relies primarily on audio to generate correlated motion. While this approach yields high fidelity results, it offers no room for choreographers to control the artistic output. To close this gap, STREAM composes conditioning signals as additive energies through energy-based cross attention. Generation becomes editable, composable, and controllable, and edits remain local and semantically addressable.

Control. I propose to move control from the output of the model into its generative process. Variance Path Diffusion (VPD) makes the noise schedule of a diffusion model an explicit random latent, a Gamma process variance path with a closed-form posterior. A Gamma-Dirichlet duality separates the total noise budget, which an observation determines from its allocation, which it does not, so the allocation becomes a design variable. Human motion supplies the groups over which that budget is spent, since limbs and successive phrases do not all require the same freedom, and music supplies the condition that allocates it, since music constrains the body at every instant in a way that text and class labels cannot. Conditioning the path on music leaves the denoiser output untouched, which makes this a control axis orthogonal by construction to conditioning and complementary to it.

Together these establish music as a structured prior that operates at every level of the generative stack. It serves as evidence for inference, as condition for generation, and as control over the generative process itself.

Bio

Seong Jong Yoo is a PhD student in the Department of Computer Science at the University of Maryland, College Park, advised by Dr. Cornelia Fermüller and Prof. Yiannis Aloimonos. His research studies human motion in musical settings, covering both the inference of motion from audiovisual recordings of performance and the generation of motion conditioned on music and language. He received the B.S. and M.S. degrees in mechanical engineering from Soongsil University, South Korea, in 2017 and 2019, respectively. His master's work was on data-driven state estimation for magnetic bearings. He then served three years of alternative military service at the Korea Institute of Science and Technology, where he worked on 3D shape reassembly. His work has published at ICCV, ICRA, WACV, and ECCV.

This talk is organized by Migo Gui