IMFD: End-to-end Multi-Face Forgery Detection through Instruction-based Large Vision-Language Models
Authors & Institutions
Dasom Choi
Chungnam National University
Sangjun Moon
Chungnam National University
Hyeongchan Im
Chungnam National University
Jaeeon Park
Institute of Science Tokyo
Jingun Kwon
Chungnam National University
Hidetaka Kamigaito
Nara Institute of Science and Technology
Taro Watanabe
Nara Institute of Science and Technology
Manabu Okumura
Institute of Science Tokyo
What Problem It Solves
IMFD reformulates multi-face forgery detection as an instruction-grounded vision-language task that jointly localizes faces and assigns per-face authenticity labels. It tests whether explicit face coordinates and image context help a large vision-language model align left-to-right face indices with the correct outputs.
Key Result
With ground-truth coordinates on the full test set, IMFD reaches 0.98 micro-F1, 0.98 macro-F1, and 0.94 exact-match accuracy; the subset with more faces scores 0.97, 0.98, and 0.87. In the true single-stage setting, those full-set metrics fall to 0.90, 0.84, and 0.77, exposing localization as the remaining bottleneck. Face-box AP is 81.9 on the full set and 83.6 on the many-face subset. Raising input resolution from 112 to 672 pixels increases controlled micro-F1 from 0.917 to 0.981 and exact match from 0.654 to 0.948. IMFD outperforms the compared baselines but is slower than the fastest single-stage method and still relies on synthetic benchmark manipulations.
Abstract
The rapid increase of deepfakes has raised significant concerns due to their spread on social media. Traditional multi-face forgery detectors crop and verify each face independently, ignoring background context and inter-face relationships, which often yields suboptimal performance. To overcome these limitations, we leverage instruction-based Large Vision-Language Models (LVLMs), which can interpret entire images and follow complex textual instructions. We propose a simple yet effective single-stage multi-face forgery detector, called IMFD (Instruction-based Multi-face Forgery Detector), which is trained end-to-end to jointly localize faces and predict per-face forgery labels. Rather than treating face box prediction only as a joint objective, IMFD explicitly integrates predicted face bounding boxes into the instruction as visual cues that enhance instruction grounding and forgery detection. To support the training and evaluation of IMFD, we convert existing multi-face forgery datasets into an instruction-based format. Experimental results and analyses show that IMFD improves multi-face forgery detection by integrating face bounding boxes into the instruction, and consistently outperforms various state-of-the-art methods.
Research Starting Point
Multi-person photos break the usual crop-then-classify assumption. Independent face crops discard scene and inter-face context, require sequential work for every person, and propagate detector or ordering errors into the forensic decision. A useful system must say which faces are fake, not merely whether an image contains any manipulation.
Method
The system converts OpenForensics samples into instructions naming faces from left to right. A CLIP-based anchor stage scores candidate regions, fuses overlapping boxes, and injects predicted coordinates into the text instruction alongside the image representation. The LVLM then returns the forged face indices and boxes. The authors evaluate both a controlled two-stage setting using ground-truth coordinates and an end-to-end single-stage setting using IMFD's predicted boxes, reporting micro-F1, macro-F1, exact image-level match, box AP, latency, and throughput.
Paper Summary
Whole-image reasoning improves crowded-scene forensics, but the gap between supplied and predicted boxes is large. Procurement tests should score exact per-face decisions and localization failures, not only image-level AUC.