log in  |  register  |  feedback?  |  help  |  web accessibility
PhD Defense: Addressing Pitfalls of Societal-Scale Mobility Data use in Urban Computing: Privacy, Bias and contamination of systems and models
Naman Awasthi
IRB-4109
Thursday, June 18, 2026, 11:59 am-2:00 pm
  • You are subscribed to this talk through .
  • You are watching this talk through .
  • You are subscribed to this talk. (unsubscribe, watch)
  • You are watching this talk. (unwatch, subscribe)
  • You are not subscribed to this talk. (watch, subscribe)
Abstract

Mobility data from smartphones, vehicle GPS (bus, car, e-bike, trucks), and shared-mobility services offer high-resolution insight into transportation behavior, public-health dynamics and infrastructure needs. This dissertation identifies pitfalls of adopting societal-scale mobility data, systems and models in urban computing by uncovering privacy boundaries, biases and contamination of such systems and models.

First, using a nationally representative vignette survey, we measure U.S. privacy perceptions for ~450 combinations of Actors, purposes and features. The survey uncovers comfort levels for different contexts in which location features may / may not be appropriate to share. We also find broad public acceptance of academic use of third‑party mobility data for public transit planning and disease tracking. Motivated by those findings, we pursue two empirical paths: Studying public transit utilization using crowdsourced data; Using location data from data aggregators to efficiently forecast spread of COVID-19.

Second, we develop and deploy BALTO, a privacy‑respectful Android app that crowdsources fine‑grained transit trajectories and trip experience surveys in Baltimore (200+ participants, ~1,200 trips). After cleaning trajectories, filtering non‑transit activity (due to gamification), and mapping trips to exact routes (eg. Bus 51 Towson-Downtown), we compare observed travel times to route‑planning APIs (Transit, GTFS and Google Maps) to identify socio‑economic and structural drivers of longer observed trip times and derive correction formulas useful for urban planners utilizing routing API services.

Third, we analyze how incorporating aggregated mobility from third‑party vendors affects COVID‑19 forecasting across U.S. counties. We quantify biases introduced by mobility signals (associated to socio-economic characteristics of the county) and introduce DemOpts, a fair regression training framework that improves equity of county‑level forecasts when mobility features are used.

Finally, we investigate contamination using a frequentist definition of memorization in large language models applied to geospatial tasks such as trajectory prediction, footfall estimation and destination prediction. We evaluate 8 open-source large language models and demonstrate evidence of data contamination for standard trajectory, Origin-Destination and Points of Interest (POI) datasets leading to inflation of performance on open datasets compared to closed data.

Together, these contributions advance methods for responsible collection, evaluation, and fair use of mobility data and systems in societal-scale tasks.

Bio

Naman Awasthi is a PhD student advised by Prof. Vanessa Frias-Martinez in the University of Maryland, College Park. His research focus is in enabling and enhancing mobility data and systems in domains like privacy policy, public transit and public health

This talk is organized by Migo Gui