REVIEW 4 cited by
PFLD: A Practical Facial Landmark Detector
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Being accurate, efficient, and compact is essential to a facial landmark detector for practical use. To simultaneously consider the three concerns, this paper investigates a neat model with promising detection accuracy under wild environments e.g., unconstrained pose, expression, lighting, and occlusion conditions) and super real-time speed on a mobile device. More concretely, we customize an end-to-end single stage network associated with acceleration techniques. During the training phase, for each sample, rotation information is estimated for geometrically regularizing landmark localization, which is then NOT involved in the testing phase. A novel loss is designed to, besides considering the geometrical regularization, mitigate the issue of data imbalance by adjusting weights of samples to different states, such as large pose, extreme lighting, and occlusion, in the training set. Extensive experiments are conducted to demonstrate the efficacy of our design and reveal its superior performance over state-of-the-art alternatives on widely-adopted challenging benchmarks, i.e., 300W (including iBUG, LFPW, AFW, HELEN, and XM2VTS) and AFLW. Our model can be merely 2.1Mb of size and reach over 140 fps per face on a mobile phone (Qualcomm ARM 845 processor) with high precision, making it attractive for large-scale or real-time applications. We have made our practical system based on PFLD 0.25X model publicly available at \url{http://sites.google.com/view/xjguo/fld} for encouraging comparisons and improvements from the community.
Forward citations
Cited by 4 Pith papers
-
NeoLoc-68: End-to-end 68-point neonatal facial landmark localisation in neonatal clinical environments
A YOLO keypoint model trained on 37k+ public images plus 1k neonatal frames achieves SOTA NME and low failure rates for 68-point neonatal landmark detection in clinical conditions.
-
Audio-Assisted Face Video Restoration with Temporal and Identity Complementary Learning
GAVN outperforms state-of-the-art face video restoration on compression artifact removal, deblurring, and super-resolution by fusing audio, landmark, and temporal identity features.
-
ORFormer: Occlusion-Robust Transformer for Accurate Facial Landmark Detection
ORFormer uses per-patch messenger tokens to detect and recover occluded face regions, reducing landmark error on WFLW and COFW.
-
Facial Expression Recognition in the Deep Learning Era: A Systematic Multi-Criteria Review of Methods, Models, Datasets, Performance, Challenges, and Future Research Directions
This survey organizes deep learning FER literature into five evolutionary phases and a seven-criteria taxonomy, compares datasets and performance, and outlines challenges.
Discussion (0). Continue with ORCID to comment.