FakeI2V-Bench: Benchmarking the Applicability of Image-level Deepfake Detectors for Deepfake Video Detection
Authors & Institutions
Pei Li
School of Cyber Science and Technology, Shandong University, Qingdao, China
Sihan Chen
School of Cyber Science and Technology, Shandong University, Qingdao, China
Delong Ran
Institute for Network Sciences and Cyberspace, BNRist, Tsinghua University, Beijing, China
Tianshuo Cong
School of Cryptologic Science and Engineering, Shandong University, Jinan, China
What Problem It Solves
FakeI2V-Bench supplies a unified test of image- and video-level detection and introduces a reusable bridge that turns frame scores into a video decision. It quantifies accuracy, cross-generator behavior, latency, model size, and performance when only 20% of a clip is manipulated.
Key Result
With the best naive aggregation, the strongest image detector reaches 80.16% AUC, already edging the top video model FTCN at 79.99%. After IV-Bridge, eleven of twelve image detectors beat FTCN; RINE-IV leads at 93.80% AUC and 97.63% AP, while the 4.34-million-parameter Patch-IV reaches 89.26% AUC and 96.12% AP. Across fourteen generators, enhanced image models average 95.51% AUC versus 91.01% for the best video detector, although Sora remains difficult at 72.76% mean AUC. With only 20% forged frames, every model drops and Patch-IV leads at 77.15% AUC.
Abstract
Recent advances in video generation models have significantly intensified the deepfake threat, yet the current deepfake video detection benchmarks remain underdeveloped. In particular, the effectiveness of image-level detectors in the video domain has not been systematically assessed. To fill this gap, we present FakeI2V-Bench, a benchmark for evaluating state-of-the-art video-level deepfake detectors in challenging scenarios, with a particular focus on systematically assessing the performance of image-level deepfake detectors in the video domain. FakeI2V-Bench comprises 97,548 videos, containing content generated by the latest powerful generation models and covering a broader range of categories. Using this dataset, we conduct a systematic evaluation of eight video-level detectors and twelve representative image-level detectors. Experimental results show that the best-performing image-level detector achieves an 80.16% AUC, slightly outperforming the strongest video-level detector (i.e., 79.99% AUC). Going beyond benchmarking, we present IV-Bridge, a general framework that enhances the applicability of image-level deepfake detectors to videos. IV-Bridge employs a random forest model with statistical features to aggregate frame-level predictions, allowing eleven image-level detectors to surpass state-of-the-art video-level approaches, with the best-performing variant achieving a 93.80% AUC. Overall, FakeI2V-Bench establishes a rigorous benchmark for deepfake video detection and introduces a novel pathway for extending image-level detectors to the video domain, offering new insights and directions for future research. Code and data are available at https://github.com/CryptoAILab/FakeI2V-Bench.
Research Starting Point
Image deepfake detectors benefit from mature training pipelines, but video benchmarks usually compare only video-specific architectures and apply frame models with simplistic averaging. That leaves teams unsure whether temporal networks are necessary, how recent video generators change the ranking, and whether an existing image detector can be adapted economically to video. Sparse manipulations, where only a fraction of frames are fake, make naive aggregation especially risky.
Method
The benchmark combines 97,548 videos from four facial and general-generation datasets and evaluates eight video detectors plus twelve representative image detectors. The authors first compare fixed score aggregations such as mean, maximum, and minimum. IV-Bridge then performs video-frame fine-tuning on non-overlapping FF++ and GenVideo frames and feeds statistical descriptors of frame-level scores to a random forest with as many as 800 trees and depth eight. Results are broken down by fourteen generation models, facial versus general content, deployment cost, and short-term forgery.
Paper Summary
An existing image detector can become a competitive video service if it is tuned on real video frames and paired with learned temporal score aggregation. The bridge is attractive for lightweight deployments, but sparse edits remain a clear failure mode and the largest accuracy winner is not the cheapest model.