log in  |  register  |  feedback?  |  help  |  web accessibility
PhD Defense: Measuring What Matters in Trustworthy AI: Certified Robustness to Agentic Safety
Shoumik Saha
IRB-3137 https://umd.zoom.us/j/8566872628?pwd=VDJ1WWZCamE2Ym9ZcGh2RjZ6YVY1Zz09&omn=99899553792&jst=2
Friday, July 24, 2026, 11:30 am-1:00 pm
  • You are subscribed to this talk through .
  • You are watching this talk through .
  • You are subscribed to this talk. (unsubscribe, watch)
  • You are watching this talk. (unwatch, subscribe)
  • You are not subscribed to this talk. (watch, subscribe)
Abstract

As AI systems are deployed in increasingly security-critical settings, their trustworthiness must be evaluated under the conditions that shape real-world risk, including adversarial manipulation, ambiguous and mixed-provenance inputs, multi-step tool use, and failures with consequential downstream effects. Establishing trust, therefore, requires moving beyond clean accuracy and isolated model responses toward guarantees and evaluations that capture robustness, reliability, executable behavior, and vulnerabilities across the broader AI agent ecosystem. My work advances a measurement-first agenda for trustworthy AI that combines two complementary approaches: certified robustness for well-specified threat models, and realistic adversarial evaluation for complex generative and agentic systems where formal guarantees are incomplete.

On the certified side, DRSM develops a de-randomized smoothing methodology for malware detection that provides formal robustness certificates against bounded byte-level perturbations while maintaining competitive standard accuracy. This work demonstrates that security-critical classifiers can move beyond empirical robustness toward provable guarantees, and contributes a public dataset of recent benign executables to support reproducible evaluation.

For generative AI, my works develop benchmarks and adversarial methods that expose failure modes overlooked by conventional evaluations. APT-Eval studies AI-polished writing, where human-authored text is minimally refined by language models, and shows that existing AI-text detectors frequently produce high false-positive rates, fail to distinguish different levels of AI involvement, and exhibit sensitivity to the choice of polishing model and writing domain. BEAST introduces an efficient, gradient-free beam-search attack that enables fast adversarial stress testing of language models under limited computational budgets. It demonstrates that aligned models can be efficiently manipulated to generate harmful responses, produce more hallucinated or irrelevant outputs, and become more vulnerable to membership-inference attacks.

Building on these model-level evaluations, my later works address agentic safety, where risk is determined not only by generated text but also by planning, tool use, workspace context, and downstream execution. JAWS-Bench evaluates code agents across empty, single-file, and multi-file workspaces, together with a hierarchical judge framework that measures compliance, harmfulness, syntactic correctness, and runtime executability. The results show that agentic interaction can overturn initial refusals and convert harmful intent into deployable code, motivating execution-aware safety evaluation. Finally, my work on semantic supply-chain attacks studies vulnerabilities in AI agent skill registries. By modifying only the natural-language content of skill descriptions, an attacker can manipulate skill discovery, bias agent selection, and evade registry governance without substantially changing the underlying functionality.

Together, these works develop a unified framework for measuring trustworthy AI across increasingly complex systems – from certifiably robust classifiers, to generative models under adversarial prompting, to tool-using agents and the ecosystems from which they acquire capabilities. The central goal is to ensure that evaluation remains aligned with the level at which real-world harm occurs: not only in model predictions or isolated responses, but also in trajectories, executable outcomes, and external capability supply chains.

Bio

Shoumik Saha is a fourth-year Ph.D. candidate in Computer Science at the University of Maryland, College Park, advised by Professor Soheil Feizi. His research focuses on the reliability, security, and trustworthiness of generative AI systems, including LLM alignment, hallucination mitigation, adversarial attacks and defenses, AI-generated text detection, and AI-agent safety. He has completed two research internships at Amazon Web Services. His work has been published at leading venues, including ICLR, ICML, NeurIPS, ACL, EMNLP, TACL, and SaTML. His research has also received media coverage from outlets such as The New York Times and The Register.

This talk is organized by Migo Gui