For almost the past decade, face recognition has been dominated by convolutional neural networks. Margin-based losses (CosFace, ArcFace) have achieved over 99.6% verification accuracy on classic benchmarks. For a long time, further performance gains were believed to rely solely on larger-scale datasets and more sophisticated loss functions. This perspective has since shifted. Having demonstrated strong performance across general computer vision tasks, Vision Transformers are now adopted for face recognition to directly remedy key CNN deficiencies. This paper explores the architectural evolution of margin‑based CNNs, wherein hybrid CNN‑Transformer designs and fully Transformer‑based systems have emerged, drawing upon a range of relevant studies published between 2024 and 2026, such as LVFace (ICCV 2025), NPTFace and PaCo-FR (CVPR 2026), and the FaceLiVT series for mobile deployment, etc. The findings reveal that the the core issue lies not in Transformers outperforming CNNs on existing benchmarks, but in the shifts in computational models underlying face‑representation mechanisms.
Research Article
Open Access