Using AI well can take expertise, context, and judgment that users may not be equipped with. Standard NLP evaluation remains output-centered, asking whether a model response satisfies a quality metric, and not whether users can apply that response situationally in-context. In this proposal, we focus on tasks with verification barriers, which presents settings where a user may have difficulty readily judging whether outputs are correct, useful, or applicable to their goal. We examine two sources of barriers: language and domain, across several case studies, culminating in the proposed work where both sources may compound for lay-users.
This proposal advances a three-stage argument. Part I diagnoses why generic explanation-centered support does not reliably help people overcome verification barriers, as reasoning traces may contain pedantic or outright erroneous explanations for an output, and generic natural language explanations can increase confidence and overreliance without improving decision accuracy. Part II studies how researchers make sense of AI-mediated information in open-ended knowledge work, deriving two design patterns: allow AI to reduce costs of retrieving relevant information while centering human judgment in refinement, and support users in distinguishing differences between source and AI-generated inferences. Part III tests whether interfaces implementing these design patterns improve appropriate reliance in unemployment guidance, where verification barriers may compound.
Together, this proposal contributes evaluation methods, system designs, and mixed-methods studies that advance how to design better human-centered NLP systems by examining how users interpret and appropriately rely on AI systems that mediate complex or unfamiliar information, especially when users lack the language or domain expertise needed to verify their outputs.
lvin Bao is a PhD student in Computer Science at the University of Maryland, advised by Marine Carpuat and Joel Chan. His research focuses on human-centered natural language processing, spanning multilingual NLP, explainable AI, and computational support for sensemaking.

