FreqFLD: Towards All-in-One Facial Landmark Detection via Frequency Modulation
Authors & Institutions
Shun Ren
College of Computer and Information Technology, China Three Gorges University
Kaijie Jin
College of Computer and Information Technology, China Three Gorges University
Shengkai Hu
School of Information Engineering, Zhongnan University of Economics and Law
Beihang Song
National Institute of Natural Hazards, Ministry of Emergency Management of China
Hang Sun
College of Computer and Information Technology, China Three Gorges University
Wenwen Min
School of Information Science and Engineering, Yunnan University
Youfa Liu
School of Computer Science, Wuhan University
Jun Wan
School of Information Engineering, Zhongnan University of Economics and Law
School of Computer Science and Engineering, Nanyang Technological University
What Problem It Solves
FreqFLD targets one all-in-one landmark detector that can be trained jointly on 300W, AFLW, COFW, and WFLW while remaining competitive on each benchmark and robust on their difficult subsets. The paper asks whether explicit frequency priors can guide expert specialization more reliably than unconstrained mixture-of-experts routing.
Key Result
The unified model reaches 2.72 NME on 300W Common, 4.52 on 300W Challenging, and 4.50 on the WFLW test set; WFLW pose, illumination, occlusion, and blur subsets score 7.54, 4.55, 5.53, and 5.15. On heavily occluded COFW it reports 4.80 NME and a 0.39% failure rate. The full frequency modules and routing loss improve the 300W challenging baseline from 4.71 to 4.52, while adding all four training datasets improves a single-dataset 4.97 result to 4.52. Performance is competitive rather than uniformly best, including 1.69 NME on AFLW where older specialized methods report lower error.
Abstract
Recent progress in deep learning has significantly advanced facial landmark detection. However, most existing methods process features in a spatial-domain manner under a dataset-specific training paradigm, which overlooks the fact that facial landmark detection is inherently geometry-driven and sensitive to frequency variations, thereby limiting cross-dataset generalization under complex scenarios and hindering the development of a facial landmark detection model. To address this issue, we propose \textbf{FreqFLD}, a \textbf{freq}uency-modulated framework towards All-in-One \textbf{f}acial \textbf{l}andmark \textbf{d}etection. Specifically, FreqFLD introduces a Frequency Modulation Module (FreqMoM) to explicitly induce the frequency prior by decoupling and modulating low- and high-frequency components, which is then injected into subsequent feature modeling to enable balanced modeling of global facial structure and local landmark details. Furthermore, FreqFLD employs a Frequency-Modulated Mixture-of-Experts (FreqMoE), with expert selection adaptively conditioned on frequency-modulated priors, enabling flexible modeling of heterogeneous facial landmark patterns under diverse and challenging scenarios. To regularize frequency-consistent modeling under the All-in-One paradigm, we further introduce a Frequency-Consistent Routing (FreqCR) loss, which constrains the routing and assignment of frequency-aware experts to promote balanced expert utilization across diverse facial scenarios, thereby enabling stable expert specialization and achieving robust facial landmark detection. Extensive experiments demonstrate that the proposed FreqFLD achieves comparable performance on popular datasets. The code is available at: https://github.com/jkj1059657014/FreqFLD.
Research Starting Point
Facial landmark models are usually trained per dataset even though production inputs mix landmark conventions, poses, occlusions, blur, and illumination. Pure spatial features can overfit one annotation regime and miss the distinction between low-frequency face structure and high-frequency landmark detail. A unified model also creates routing conflicts when one expert is expected to handle every geometry and image-quality regime.
Method
The Frequency Modulation Module decomposes features into low- and high-frequency components, models them separately, and injects the resulting prior into downstream layers. A frequency-modulated mixture of experts uses that prior to route heterogeneous facial patterns, while Frequency-Consistent Routing loss balances utilization and encourages stable specialization. Joint training equalizes the four source datasets to 20,000 samples each, yielding an 80,000-image training pool evaluated under each dataset's native landmark and normalization protocol.
Paper Summary
Frequency-aware routing makes a credible case for consolidating several landmark models into one, but teams should compare per-camera and per-landmark-scheme error because unified training does not win every individual benchmark.