IJCB-AFMFR 2026: Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data
Authors & Institutions
Tahar Chettaoui
Fraunhofer Institute for Computer Graphics Research IGD, Germany
Guray Ozgur
Fraunhofer Institute for Computer Graphics Research IGD, Germany
Technical University of Darmstadt, Germany
Eduarda Caldeira
Fraunhofer Institute for Computer Graphics Research IGD, Germany
Technical University of Darmstadt, Germany
Arturas Nakvosas
Vilnius University, Lithuania
Hatef Otroshi Shahreza
Idiap Research Institute, Switzerland
Sébastien Marcel
Idiap Research Institute, Switzerland
Rishabh Shukla
Indian Institute of Technology Jammu, India
Aditya Takkar
Indian Institute of Technology Jammu, India
Rushil Khullar
Indian Institute of Technology Jammu, India
Lalak Yadav
Indian Institute of Technology Jammu, India
Gourav Gupta
ArogyaPandit Private Limited, India
Anant Gupta
ArogyaPandit Private Limited, India
Shiqi Yu
Southern University of Science and Technology, China
Vitomir Struc
University of Ljubljana, Slovenia
Naser Damer
Fraunhofer Institute for Computer Graphics Research IGD, Germany
Technical University of Darmstadt, Germany
Fadi Boutros
Fraunhofer Institute for Computer Graphics Research IGD, Germany
What Problem It Solves
Earlier synthetic-data challenges mixed together generator quality, backbone choice, and training recipes. This benchmark fixes CLIP ViT-L/14 and the IDPERTURB generator so that full fine-tuning, LoRA, rank-stabilized LoRA, and projection-only adaptation can be compared under the same identity data and evaluation suite.
Key Result
Full fine-tuning with Sub-Center ArcFace leads the Full Data Track at 95.51% average accuracy on the small benchmarks and 87.42% TAR at FAR 1e-5 on IJB-C, versus 76.71% for the FRoundation baseline. Under limited data, rsLoRA is more robust: Idiap-BSP reaches 94.52% average accuracy and 81.13% IJB-C TAR at FAR 1e-5. The fairness evaluation also changes by regime, with the full-data winner reaching 91.70% average RFW accuracy and the limited-data winner 88.92%.
Abstract
This paper presents a summary of the Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data (AFMFR), held at the 2026 International Joint Conference on Biometrics (IJCB 2026). The competition received a total of eight valid submissions from four distinct teams across two complementary tracks: a Full Data Track, in which participants adapt the CLIP ViT-L/14 foundation model using large-scale synthetic identity data, and a Limited Data Track, designed to reflect more resource-constrained adaptation regimes. All training data was generated exclusively using IDPERTURB. Submitted solutions are ranked based on verification and identification performance across a diverse suite of benchmarks, including LFW, CFP-FP, AgeDB-30, CALFW, CPLFW, IJB-B, IJB-C, and TinyFace, using the Borda count method. Fairness evaluation is additionally conducted on the RFW dataset across four demographic groups. The results demonstrate that adaptation of the CLIP foundation model with synthetic training data substantially improves over the off-the-shelf model and, in several cases, surpasses the baseline. Notably, full fine-tuning with Sub-Center ArcFace (DMSTI-Neurotechnology) leads the Full Data Track, while rank-stabilized LoRA adaptation (Idiap-BSP) proves most effective under limited-data conditions.
Research Starting Point
Face recognition developers increasingly lack broad, consented real-face training sets, while general vision foundation models are not sufficiently discriminative at strict biometric operating points without adaptation. The competition asks whether synthetic identities can provide that adaptation and how much trainable capacity is appropriate when the available synthetic set changes by more than an order of magnitude.
Method
The competition defines a Full Data Track with about three million images from 35,000 synthetic identities and a Limited Data Track with 250,000 images from 5,000 identities. Eight valid submissions are scored with a Borda aggregation across five small verification sets, IJB-B, IJB-C, and TinyFace, while RFW measures performance dispersion across four demographic groups. The submitted methods explore full-backbone updates, LoRA and rsLoRA ranks, margin losses, augmentation, and lightweight projection tuning.
Paper Summary
For teams adapting a large visual backbone, synthetic data volume should drive the tuning strategy. Full fine-tuning pays off at multi-million-image scale, while high-rank rsLoRA acts as a useful regularizer when identities and images are constrained; a single recipe is unlikely to be optimal for both settings.