← Back to Blog
Research RadarFace DetectionFacial LandmarksarXivSeptember 2026

Monthly arXiv Radar

September 2026 Face Detection Papers: Unified Landmarks, Diffusion Alignment, and Classroom Deployment

Direct face-box papers were sparse in September, so this roundup follows the adjacent localization stack that production systems need after a box is found. FreqFLD trains one frequency-aware landmark model across four datasets. FAHCD-Net treats landmark heatmaps as a conditional denoising problem under pose, occlusion, blur, and illumination changes. The Visage study brings detection and recognition back together in crowded classroom imagery, where throughput and missed faces matter as much as top-line identity accuracy.

What This Month Signals

FreqFLD balances global structure and local detail by splitting frequency components and routing samples through frequency-conditioned experts; one jointly trained model reaches 4.50 NME on WFLW and 4.80 NME with a 0.39% failure rate on occluded COFW. FAHCD-Net instead cascades conditional diffusion stages and suppresses redundant high-frequency heatmap noise, reducing 300W challenging-set NME from 4.48 at stage one to 4.29 at stage three. The Visage study reports that a 3.2-million-parameter YOLOv8n reaches 0.935 mAP@0.5, while much larger detectors gain accuracy at a substantial parameter cost; several recognizers reach 99.75% Top-1 on its 100-class recognition set. These results are promising, but the two landmark papers remain benchmark studies and the classroom data comes from a limited institutional setting, so cross-camera validation is still required.

Paper 012026-09-09cs.CV

FreqFLD: Towards All-in-One Facial Landmark Detection via Frequency Modulation

Authors & Institutions

Shun Ren

College of Computer and Information Technology, China Three Gorges University

Kaijie Jin

College of Computer and Information Technology, China Three Gorges University

Shengkai Hu

School of Information Engineering, Zhongnan University of Economics and Law

Beihang Song

National Institute of Natural Hazards, Ministry of Emergency Management of China

Hang Sun

College of Computer and Information Technology, China Three Gorges University

Wenwen Min

School of Information Science and Engineering, Yunnan University

Youfa Liu

School of Computer Science, Wuhan University

Jun Wan

School of Information Engineering, Zhongnan University of Economics and Law

School of Computer Science and Engineering, Nanyang Technological University

What Problem It Solves

FreqFLD targets one all-in-one landmark detector that can be trained jointly on 300W, AFLW, COFW, and WFLW while remaining competitive on each benchmark and robust on their difficult subsets. The paper asks whether explicit frequency priors can guide expert specialization more reliably than unconstrained mixture-of-experts routing.

Key Result

The unified model reaches 2.72 NME on 300W Common, 4.52 on 300W Challenging, and 4.50 on the WFLW test set; WFLW pose, illumination, occlusion, and blur subsets score 7.54, 4.55, 5.53, and 5.15. On heavily occluded COFW it reports 4.80 NME and a 0.39% failure rate. The full frequency modules and routing loss improve the 300W challenging baseline from 4.71 to 4.52, while adding all four training datasets improves a single-dataset 4.97 result to 4.52. Performance is competitive rather than uniformly best, including 1.69 NME on AFLW where older specialized methods report lower error.

Abstract

Recent progress in deep learning has significantly advanced facial landmark detection. However, most existing methods process features in a spatial-domain manner under a dataset-specific training paradigm, which overlooks the fact that facial landmark detection is inherently geometry-driven and sensitive to frequency variations, thereby limiting cross-dataset generalization under complex scenarios and hindering the development of a facial landmark detection model. To address this issue, we propose \textbf{FreqFLD}, a \textbf{freq}uency-modulated framework towards All-in-One \textbf{f}acial \textbf{l}andmark \textbf{d}etection. Specifically, FreqFLD introduces a Frequency Modulation Module (FreqMoM) to explicitly induce the frequency prior by decoupling and modulating low- and high-frequency components, which is then injected into subsequent feature modeling to enable balanced modeling of global facial structure and local landmark details. Furthermore, FreqFLD employs a Frequency-Modulated Mixture-of-Experts (FreqMoE), with expert selection adaptively conditioned on frequency-modulated priors, enabling flexible modeling of heterogeneous facial landmark patterns under diverse and challenging scenarios. To regularize frequency-consistent modeling under the All-in-One paradigm, we further introduce a Frequency-Consistent Routing (FreqCR) loss, which constrains the routing and assignment of frequency-aware experts to promote balanced expert utilization across diverse facial scenarios, thereby enabling stable expert specialization and achieving robust facial landmark detection. Extensive experiments demonstrate that the proposed FreqFLD achieves comparable performance on popular datasets. The code is available at: https://github.com/jkj1059657014/FreqFLD.

Research Starting Point

Facial landmark models are usually trained per dataset even though production inputs mix landmark conventions, poses, occlusions, blur, and illumination. Pure spatial features can overfit one annotation regime and miss the distinction between low-frequency face structure and high-frequency landmark detail. A unified model also creates routing conflicts when one expert is expected to handle every geometry and image-quality regime.

Method

The Frequency Modulation Module decomposes features into low- and high-frequency components, models them separately, and injects the resulting prior into downstream layers. A frequency-modulated mixture of experts uses that prior to route heterogeneous facial patterns, while Frequency-Consistent Routing loss balances utilization and encourages stable specialization. Joint training equalizes the four source datasets to 20,000 samples each, yielding an 80,000-image training pool evaluated under each dataset's native landmark and normalization protocol.

Paper Summary

Frequency-aware routing makes a credible case for consolidating several landmark models into one, but teams should compare per-camera and per-landmark-scheme error because unified training does not win every individual benchmark.

Paper 022026-09-15cs.CV

FAHCD-Net: Frequency-Adaptive Heatmap-Conditional Diffusion Networks for Robust Facial Landmark Detection

Authors & Institutions

Jun Wan

School of Information Engineering, Zhongnan University of Economics and Law

Jiwei Hu

School of Information Engineering, Zhongnan University of Economics and Law

Shengkai Hu

School of Information Engineering, Zhongnan University of Economics and Law

Qilu Zhu

School of Information Engineering, Zhongnan University of Economics and Law

What Problem It Solves

FAHCD-Net seeks robust landmark localization under pose, occlusion, illumination, and blur by making a heatmap-conditioned diffusion process explicitly frequency aware. It also tests whether a cascade can progressively improve an initial mean-shape condition instead of relying on a high-quality detector-generated starting heatmap.

Key Result

FAHCD-Net reports 2.51 NME on 300W Common, 2.99 on the full set, 1.17 on AFLW frontal, and 2.08 on AFLW full. On 300W Challenging, successive diffusion stages reduce NME from 4.48 to 4.33 and then 4.29, compared with 4.81, 4.73, and 4.69 when conditioning on a mean shape. Adding COFW, WFLW, and AFLW training data improves the reported 300W challenging result from 4.76 to 4.49 in the multi-dataset ablation. The paper demonstrates benchmark robustness but does not report detector-plus-landmark latency, so the cost of iterative diffusion remains a deployment question.

Abstract

Facial Landmark Detection(FLD) is a crucial task in various applications and has achieved significant advancements in recent years. However, current FLD methods still struggle under challenging conditions, where facial structural variations, information loss, and noise interference severely compromise the integrity and accuracy of learned facial features. To address these issues, we propose Frequency-Adaptive Heatmap-Conditional Diffusion Network (FAHCD-Net), which integrates a Frequency-Adaptive Heatmap-Conditional Diffusion (FAHCD) model with a Smoothness Regularization (SR) loss in a cascaded framework. Specifically, the FAHCD model incorporates a Hierarchical Frequency Adaptation (HFA) module designed to suppress redundant high-frequency noise through multi-layer frequency decomposition and adaptive reconstruction, thereby preserving essential facial structures. Additionally, the SR loss is proposed to further mitigate the interference of high-frequency noise and enhance the smoothness of the generated landmark heatmaps. By cascading the FAHCD model with the SR loss, FAHCD-Net effectively leverages both statistical and frequency-based distribution characteristics of the data to progressively generate more accurate landmark heatmaps from noisy inputs. Extensive experiments on popular benchmarks demonstrate the effectiveness and robustness of the proposed method, achieving state-of-the-art performance in FLD tasks under challenging scenarios. The source code is available at https://github.com/HJWKryptonite/FAHCD-Net.

Research Starting Point

Landmark heatmaps deteriorate when pose or expression changes the face structure, when occlusion removes evidence, and when blur or lighting injects high-frequency noise. Standard regression treats those effects as direct prediction errors. Diffusion offers iterative recovery from noisy or incomplete evidence, but unconstrained denoising can itself produce spurious high-frequency heatmap energy and unstable peaks.

Method

Each cascade stage uses a Frequency-Adaptive Heatmap-Conditional Diffusion model to refine landmark distributions. Hierarchical Frequency Adaptation repeatedly decomposes and reconstructs features so low-frequency facial structure is retained while redundant high-frequency noise is suppressed. A smoothness regularization term aligns generated heatmap frequency behavior with the ground truth. Multi-dataset training adds COFW, WFLW, and AFLW to 300W, and three stages pass their predicted heatmaps forward as the next condition.

Paper Summary

Conditional diffusion can recover cleaner landmark heatmaps in difficult imagery, but its iterative cost should be measured against the camera's latency budget and a strong non-diffusion alignment baseline.

Paper 032026-09-19cs.CV

Towards Robust Classroom Attendance: A Comprehensive Evaluation of Face Detection and Recognition Models

Authors & Institutions

Himani Trivedi

Computer Engineering Department, LDRP Institute of Technology and Research, Kadi Sarva Vishwavidyalaya

Hiren Patel

Vidush Somany Institute of Technology and Research, Kadi Sarva Vishwavidyalaya

Ridham Patel

Information Technology Department, LDRP Institute of Technology and Research, Kadi Sarva Vishwavidyalaya

Krutika Patel

Information Technology Department, LDRP Institute of Technology and Research, Kadi Sarva Vishwavidyalaya

Nancy Patel

Information Technology Department, LDRP Institute of Technology and Research, Kadi Sarva Vishwavidyalaya

What Problem It Solves

The study introduces paired detection and recognition datasets for classroom conditions and compares a family of detector sizes with seven current face recognition architectures. Its goal is to identify a practical accuracy-efficiency operating point rather than assume the largest detector or recognizer is the best attendance component.

Key Result

YOLOv8n uses 3.2 million parameters and reaches 0.932 precision, 0.886 recall, 0.91 F1, and 0.935 mAP@0.5; larger variants approach roughly 0.97 mAP@0.5 but use 43.7 million to 68.2 million parameters. In recognition, FaceLiVTv2-L, EdgeFace Base, LVFace ViT-B, and TransFace ViT-B each reach 99.75% Top-1 in the reported table. EdgeFace Base does so in 4.029 ms, faster than FaceLiVTv2-L at 6.081 ms, while model sizes vary widely. These are closed-set results from a limited set of educational institutions, so they do not establish spoof resistance, cross-school generalization, or policy suitability for automated attendance.

Abstract

Manual attendance methods, such as paper or register-based systems, take a lot of time, can lead to errors, and are easy to falsify. Face recognition is more reliable, but it frequently struggles in classrooms because lighting and other conditions can vary. Face recognition datasets are designed for regulated environments and do not capture the actual challenges found in classrooms. To address this, a new face detection and recognition dataset, the Visage Face dataset, comprising 16,234 face samples, is proposed for the task of face detection and recognition. The photos are taken from different angles and under varying lighting conditions, with students showing a range of expressions, and some faces partly covered to reflect real-life situations. A YOLO-based system is used to detect faces and tested seven advanced face recognition models with thirteen configurations: LVFace, QCFace, FaceLiVTv2, TopoFR, EdgeFace, TransFace, and GhostFaceNets. Of these, FaceLiVTv2-M performed best, with 99.75% Top-1/Top-5 accuracy and an inference time of 6.459 ms. These results show that the Visage Face Dataset is a realistic and challenging benchmark for face recognition in classroom attendance.

Research Starting Point

Classroom attendance combines many small faces, varying pose and lighting, partial occlusion, and a need to process repeated video frames on affordable hardware. General face datasets do not reproduce that operating environment, while reporting recognition accuracy alone hides missed detections and computational cost. Institutions also need data collected with consent rather than repurposed identity imagery.

Method

The authors collect 801 classroom images, manually annotate faces, and use rotation, scale, translation, noise, and flipping to build the detection set; the paper reports 3,635 detection images and 96,409 face annotations after augmentation. The recognition set contains 3,500 images across 100 classes, aligned from five RetinaFace landmarks and normalized to 112 by 112. YOLOv8 n/s/m/l/x variants are evaluated for detection, and thirteen configurations of LVFace, QCFace, FaceLiVTv2, TopoFR, EdgeFace, TransFace, and GhostFaceNets are compared for recognition.

Paper Summary

The useful result is the measured detector-recognizer trade-off, not the near-perfect closed-set score. A pilot should reproduce recall, latency, consent, spoofing, and unknown-person behavior on the institution's own cameras.