Preprint 2 Mentions
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination
Shuo Liang2026
Yixing MaPengfei Zhou
Low Citations
0 citations · Artificial Intelligence
Open Access

TLDR

Computer-made videos of wars and disasters can look real enough to fool people. Current checking tools often miss them, especially after the videos are shared online.

Summary

1 Study Aim

The authors aim to test how well detectors (tools that identify computer-made videos) recognize realistic crisis footage. They examine detector performance across video-making conditions, human judgments, and social dissemination (sharing through online networks). The study also asks whether existing methods remain reliable as video generators change. The study tests whether today’s checking tools can reliably spot realistic computer-made crisis videos.

2 Study Design

The researchers created RA-Bench, a benchmark (standard testing collection) containing 17,886 videos. It includes 1,830 real-video anchors across 10 social-risk categories and 16,056 clips from four open-source and five closed-source generators. They tested seven traditional detectors, 10 zero-shot multimodal models (systems used without task-specific training), and two MLLMs, or multimodal large language models, trained for video detection. They also assessed generation quality, conditioning information, sampling seeds (random starting settings), human authenticity judgments, and social dissemination. The researchers tested many checking tools against real and computer-made crisis videos, including videos changed and shared in different ways.

3 Findings

The study reveals that none of the three detector families generalizes consistently across RA-Bench. Generation properties affect detector families differently, while source-level detection patterns remain stable across sampling seeds. The authors find that videos misleading people are also difficult for detectors to identify. Social dissemination makes detection harder, reducing detector reliability after sharing. The findings demonstrate that current methods struggle with realistic AI-generated videos and support developing detectors robust to changing generators. Existing checking tools often fail on convincing computer-made videos, particularly after people share them online.

Abstract

Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-world crises, creating substantial risks of misinformation. Existing benchmarks, however, provide limited evidence on detector and generator behavior in such settings, including how detectability varies with generation conditions, how people perceive generated videos, and whether detectors remain reliable during social dissemination. To address this gap, we introduce RA-Bench, a benchmark for AI-generated video detection that uses Real videos as Anchors. RA-Bench contains 17,886 videos, comprising 1,830 real-video anchors across 10 social-risk categories and 16,056 generated clips from four open-source and five closed-source generators. Based on RA-Bench, we organize our evaluation along three dimensions. We first assess detector generalization across seven traditional detectors, ten zero-shot multimodal models under three review settings, and two MLLMs specifically fine-tuned on AI-generated video detection. Across these methods, none of the three detector families generalizes consistently across RA-Bench instances. We then examine how detectability varies with generation quality, conditioning information, and sampling seeds. These analyses show that generation properties affect detector families differently, while source-level detection patterns remain stable across seeds. Finally, we study human authenticity judgments and detector reliability during social dissemination. We find that videos that mislead people are also difficult for current detectors, and that social dissemination makes detection harder. Together, these findings show that current methods struggle to detect realistic AI-generated videos, highlighting the need for detectors robust to evolving video generators.

Referenced In