← Back to Blog
Research RadarDeepfake DetectionMedia ForensicsarXivSeptember 2026

Monthly arXiv Radar

September 2026 Deepfake Detection Papers: Multi-Face Reasoning, Smartphone False Alarms, and Continual Learning

The strongest September deepfake work focuses on system behavior rather than another closed-set score. IMFD asks a vision-language model to localize and classify every face in one instruction-grounded pass. LAION-Mobile audits original detector checkpoints on modern synthetic content and nearly one million smartphone-photo records, revealing severe calibration drift. MSFD treats video detector updates as a continual-learning problem and preserves spatial, temporal, and joint frequency evidence separately.

What This Month Signals

IMFD's coordinate-aware instruction raises its two-stage result on OpenForensics to 0.98 micro-F1, 0.98 macro-F1, and 0.94 exact-match accuracy; the end-to-end single-stage version drops to 0.90, 0.84, and 0.77 because face localization still propagates errors. LAION-Mobile finds that none of twelve published detectors exceeds 0.624 AUC on a modern AI mixture and that modern-calibrated thresholds flag 17-91% of authentic phone photos; its corpus, however, barely covers current flagship neural ISPs. MSFD reaches 93.02% average AUC with 2.29 points of forgetting across six sequential video datasets, and 83.65% average accuracy with negative forgetting in a few-shot protocol. Together, the papers argue for whole-scene evaluation, current-device negatives, explicit calibration ownership, and regression testing after every detector update.

Paper 012026-09-17cs.CV

IMFD: End-to-end Multi-Face Forgery Detection through Instruction-based Large Vision-Language Models

Authors & Institutions

Dasom Choi

Chungnam National University

Sangjun Moon

Chungnam National University

Hyeongchan Im

Chungnam National University

Jaeeon Park

Institute of Science Tokyo

Jingun Kwon

Chungnam National University

Hidetaka Kamigaito

Nara Institute of Science and Technology

Taro Watanabe

Nara Institute of Science and Technology

Manabu Okumura

Institute of Science Tokyo

What Problem It Solves

IMFD reformulates multi-face forgery detection as an instruction-grounded vision-language task that jointly localizes faces and assigns per-face authenticity labels. It tests whether explicit face coordinates and image context help a large vision-language model align left-to-right face indices with the correct outputs.

Key Result

With ground-truth coordinates on the full test set, IMFD reaches 0.98 micro-F1, 0.98 macro-F1, and 0.94 exact-match accuracy; the subset with more faces scores 0.97, 0.98, and 0.87. In the true single-stage setting, those full-set metrics fall to 0.90, 0.84, and 0.77, exposing localization as the remaining bottleneck. Face-box AP is 81.9 on the full set and 83.6 on the many-face subset. Raising input resolution from 112 to 672 pixels increases controlled micro-F1 from 0.917 to 0.981 and exact match from 0.654 to 0.948. IMFD outperforms the compared baselines but is slower than the fastest single-stage method and still relies on synthetic benchmark manipulations.

Abstract

The rapid increase of deepfakes has raised significant concerns due to their spread on social media. Traditional multi-face forgery detectors crop and verify each face independently, ignoring background context and inter-face relationships, which often yields suboptimal performance. To overcome these limitations, we leverage instruction-based Large Vision-Language Models (LVLMs), which can interpret entire images and follow complex textual instructions. We propose a simple yet effective single-stage multi-face forgery detector, called IMFD (Instruction-based Multi-face Forgery Detector), which is trained end-to-end to jointly localize faces and predict per-face forgery labels. Rather than treating face box prediction only as a joint objective, IMFD explicitly integrates predicted face bounding boxes into the instruction as visual cues that enhance instruction grounding and forgery detection. To support the training and evaluation of IMFD, we convert existing multi-face forgery datasets into an instruction-based format. Experimental results and analyses show that IMFD improves multi-face forgery detection by integrating face bounding boxes into the instruction, and consistently outperforms various state-of-the-art methods.

Research Starting Point

Multi-person photos break the usual crop-then-classify assumption. Independent face crops discard scene and inter-face context, require sequential work for every person, and propagate detector or ordering errors into the forensic decision. A useful system must say which faces are fake, not merely whether an image contains any manipulation.

Method

The system converts OpenForensics samples into instructions naming faces from left to right. A CLIP-based anchor stage scores candidate regions, fuses overlapping boxes, and injects predicted coordinates into the text instruction alongside the image representation. The LVLM then returns the forged face indices and boxes. The authors evaluate both a controlled two-stage setting using ground-truth coordinates and an end-to-end single-stage setting using IMFD's predicted boxes, reporting micro-F1, macro-F1, exact image-level match, box AP, latency, and throughput.

Paper Summary

Whole-image reasoning improves crowded-scene forensics, but the gap between supplied and predicted boxes is large. Procurement tests should score exact per-face decisions and localization failures, not only image-level AUC.

Paper 022026-09-10cs.CV

LAION-Mobile: Evaluating Deepfake Detectors On One Million Smartphone Photos

Authors & Institutions

Achim von Stryk

Stralsund University, Germany

Janis Keuper

Institute for Machine Learning and Analytics, Offenburg University, Germany

What Problem It Solves

LAION-Mobile audits whether twelve original detector checkpoints generalize to modern AI imagery and how often they false-alarm on authentic smartphone photos. It separates threshold-independent discrimination from threshold calibration and asks whether legacy GAN calibration creates a misleading deployment picture.

Key Result

On NTIRE 2026, the best detector reaches only 0.624 AUC; five of twelve score below chance, and none exceeds 13.5% true-positive rate at 5% false-positive rate. Legacy-calibrated thresholds make several systems appear usable at no more than 11% phone-photo false alarms, but modern calibration makes the same detectors flag 17-91% of authentic photos. UnivFD illustrates the shift: 0.2% becomes 39.1% on identical phone images. The corpus is dominated by older devices and contains almost no current flagships, authenticity is inferred rather than certified per image, and LAION-Mobile is real-only, so it cannot provide a two-class phone AUC.

Abstract

Most Deepfake detectors report near-perfect AUC scores on their reference benchmarks. However, a recent ICML position paper argues that these evaluations collectively neglect the impact of modern smartphone photography: the widely used on-device neural image-signal processing pipelines (like multi-sensor fusion or noise and motion-blur suppression) increasingly shift the imaging paradigm from simple lens projections towards computational photography. Hence, devices actually generate, rather than record photos. This increases the risk that deepfake detectors may flag ordinary phone photos as fake. Due to the lack of large-scale datasets containing images from modern smartphones, this hypothesis has so far only been tested in small proof-of-concept studies. The aim of this paper is to close this gap. We introduce LAION-Mobile, an open dataset containing about 1 million smartphone images with EXIF metadata distilled from re-LAION-5B. Evaluating twelve state-of-the-art deepfake detectors with their original paper checkpoints on a 9,115-image evaluation sample of this pool (DIRE on 738), we report three key findings: (i) On modern AI content no detector exceeds AUC 0.624, and five of twelve fall below chance. (ii) Real-photo false-alarm rates are an artefact of threshold calibration: thresholds fitted on legacy GAN data make several detectors look deployable (less than 11 percent FPR), yet the same detectors flag 17-91 percent of real photos once the identical criterion is refit on modern content. (iii) Consequently, no detector both beats chance on modern AI content and keeps a deployable real-photo false-alarm rate. Mirroring the device mix of web collections, the corpus probes the first neural-ISP generation (2018-2020); current flagships are essentially absent, leaving the modern-ISP regime as the open gap.

Research Starting Point

Published deepfake detectors often approach perfect AUC on their own benchmarks, while modern phones apply HDR fusion, denoising, sharpening, motion suppression, and other learned image-signal processing to authentic photos. Those computational artifacts can resemble generated-image fingerprints. Without a large real-phone corpus and a threshold calibrated on current synthetic content, a detector may look safe while flagging ordinary photographs at an unusable rate.

Method

The authors filter re-LAION-5B by smartphone EXIF and non-photo heuristics to create a 935,399-image real-photo pool with device metadata, then score a fixed 9,115-image evaluation sample, except DIRE on 738 expensive reconstructions. Twelve detectors are first checked against their original or ProGAN-style benchmarks, then evaluated on the NTIRE 2026 modern real/fake mixture. Equal-error thresholds fitted separately on legacy ProGAN-ISP and modern NTIRE are transferred to the same phone photos to expose calibration drift.

Paper Summary

A near-perfect paper AUC does not transfer to modern content, and threshold ownership can change real-photo false alarms by orders of magnitude. Every deployment needs current-device negatives, a documented calibration set, and monitored operating-point drift.

Paper 032026-09-03cs.CV

Preserving Knowledge across Space and Time for Continual Video Deepfake Detection

Authors & Institutions

Taehoon Kim

Graduate School of Artificial Intelligence, Chung-Ang University

Jongwook Choi

Graduate School of Artificial Intelligence, Chung-Ang University

Heejae Jo

Department of Advanced Imaging, Graduate School of Advanced Imaging Science, Multimedia and Film, Chung-Ang University

Byungmin Park

Graduate School of Artificial Intelligence, Chung-Ang University

Jongwon Choi

Graduate School of Artificial Intelligence, Chung-Ang University

Department of Advanced Imaging, Graduate School of Advanced Imaging Science, Multimedia and Film, Chung-Ang University

Department of Metaverse Convergence, Chung-Ang University

What Problem It Solves

MSFD targets the stability-plasticity problem specifically for video deepfake detection. It preserves spatial, temporal, and spatiotemporal forensic knowledge separately while updating the detector, and evaluates full-data, generator-incremental, and few-shot sequences rather than a single static train-test split.

Key Result

In the six-task protocol, the full system reaches 93.02% average AUC with 2.29 points of average forgetting, versus 89.84% and 6.81 for the iCaRL-style baseline. In the generator-incremental protocol it reaches 96.67% average AUC, including 91.47% on image-to-video forgeries, with 1.37 forgetting. Under only 25 real and 25 fake videos per new task, it achieves 83.65% average accuracy and -1.25 forgetting, indicating positive backward transfer; IDER forgets less at -2.10 but has slightly lower 83.47% accuracy. Ablations show all three frequency branches outperform a single undivided representation. The method still shows a forgetting increase after the predominantly Asian KoDF task, signaling unresolved demographic and domain shift.

Abstract

The continuous emergence of high-quality video deepfakes requires detectors that continually adapt to new forgery patterns, yet existing approaches, which are designed for deepfake images, fail to capture video-specific cues. Unlike deepfake images that contain only spatial artifacts, deepfake videos leave distinct evidence along both spatial and temporal axes, necessitating the separate preservation of each modality during sequential model updates. To overcome this limitation, we introduce a continual deepfake video detection framework, Modality-Specific Frequency Distillation (MSFD), that explicitly decomposes video features into spatial, temporal, and spatiotemporal modalities in the frequency domain. This decomposition enables independent preservation of each modality, as different deepfake video types exhibit varying reliance on spatial and temporal cues across tasks. Furthermore, MSFD adopts a cross-modality decorrelation loss that encourages spatiotemporal representations to remain orthogonal to single-modality cues. Extensive experiments show that our framework achieves stronger adaptation and preserves performance more effectively than state-of-the-art methods across diverse continual deepfake video scenarios.

Research Starting Point

Video forgery types arrive sequentially after deployment, so a detector must learn new generators without erasing old evidence. Continual methods inherited from image classification tend to preserve one entangled representation, even though video deepfakes leave different mixtures of spatial boundaries, temporal inconsistencies, and joint space-time artifacts. A low forgetting score is also not useful if the system refuses to adapt to the new task.

Method

Intermediate video features are transformed into 2D spatial, 1D temporal, and joint spatiotemporal frequency representations. Modality-Specific Frequency Distillation aligns each spectrum with the previous model during sequential updates. A modality-specific adaptor reweights channels and learns Gumbel-Softmax masks for task-relevant frequency bands, while a decorrelation loss prevents the joint branch from duplicating the single-modality cues. Replay examples and the same 3D ResNet-18 backbone are used for most baseline comparisons across six deepfake video datasets.

Paper Summary

Separating spatial and temporal forensic memory improves detector updates, especially with little new data. The remaining KoDF shift is a reminder to regression-test both generator families and population domains after every continual-learning release.