Multimodal AI and Transportation Reasoning
How can AI reason about traffic scenes using evidence from different perspectives?
We study how vision-language models interpret roadway video, connect vehicle and infrastructure views, and reason about dynamic events. We build benchmarks and learning methods to evaluate what models observe, infer, and generalize in transportation settings.
Research directions
- Benchmarking visual understanding, temporal reasoning, and crash-scene interpretation.
- Learning from vehicle, infrastructure, and cooperative views with multimodal models.
- Developing compact models for roadside deployment and local workflows for crash narratives.
Selected work
CrashSight: A Phase-Aware, Infrastructure-Centric Video Benchmark for Traffic Crash Scene Understanding and Reasoning
CVPR Workshops · 2026
An Agentic Workflow for Detecting Personally Identifiable Information in Crash Narratives
Preprint · 2026